Showing posts with label google. Show all posts
Showing posts with label google. Show all posts

Saturday, 21 February 2015

Ask A Foolish Question

From Project Gutenberg: Robert Sheckley's 1953 short SF story Ask A Foolish Question, which is still worth reading, as well as having quite a bit of modern pertinence. The story concerns various species, including humans, who make pilgrimage to a device called Answerer, "Because Answerer knows everything".

Thursday, 13 February 2014

Ibong Adarna: Google Mistranslate

from Project Gutenberg Tagalog edition
Ibong Adarna is the title of a massively popular epic fantasy in the mythology and culture of the Philippines; it originally went under the snappy title of Corrido ng Pinagdaanang Buhay nang Tatlong Principeng, Magcacapatid na Anac nang haring Fernando at nang Reina Valeriana sa Caharian ng Berbania ("Corrido of the Traveled/Travailed Life of Three Princes, Sibling Children of King Fernando and Queen Valeriana of the Kingdom of Berbania"). Despite the Spanish names, it evidently pre-dates the Spanish Era in the Philippines.
One of the most beautiful tales which the Filipinos are wont to hear in their youth since time immemorial is the “Ibong Adarna”. This tale, or awit, is known all over the Philippines and was told vocally probably  centuries before it was anonymously printed in Tagalog in about 1860, thereafter appearing in the different vernacular dialects  — Visayan, Pampango, Ilocano and Bicol where its version varies. It is a four-thousand-one-hundred-thirty-six-line metrical romance in quartillas, of iambic tetrameter, on the life and adventures of the three sons of King Ferdinand of Berbania — one if not the most interesting of the fantastic tales in Philippine literature. According to the more reliable studies on the subject the tale is of Pre-Spanish origin and so is indigenous, although it is not free in its modern version from outside influence, like the other native corridos that were "derived" from European romances, that are greatly saturated with the "medieval flavor and setting of chivalry". It is comparable or possibly on a par with the world-famous Arabian Nights' Entertainments — a book included in the outside reading texts of both public and private schools. Although its language is not as literary as Florante at Laura, the work nevertheless indicates that it is the product of a pen of the stamp of a Balagtas, which in spite of not having the academic preparation of that prince of Tagalog poets, in fact it is like an uncut diamond which though it does not glitter as much as the cut and polished one, yet does not on that account cease to be a diamond.
- The Adarna Bird (A Filipino Tale of Pre-Spanish Origin Incorporated in the Development of Philippine Literature, the Rapid Growth of Vernacular Belles-letters from Its Earliest Inception to the Present Day), Eulogio Balan Rodriguez, General Printing Press, 1933.
See the paper ANG MGA INAGDAANANG BUHAY NG IBONG ADARNA: Narrative and Ideology in the Adarna's Corrido and Filmic Versions (Francisco Benitez, Department of Comparative Literature, University of Washington) for a detailed analysis of the story's content and cultural significance.

As you can gather from the synopsis at the Wikipedia page Ibong Adarna (mythology), it's a fairly convoluted story of the adventures of three princes, Pedro, Diego and Juan. The first two conspire against Juan, as they go on a quest to heal their ailing father (who has got ill from worrying about a dream in which two traitors conspire against Juan). The title refers to the Adarna bird that's central to the story and has properties probably unique  for a mythical creature. It has powers of magic and healing, but a dangerous aspect: petrifying poo! At the end of the day it sings seven songs that lull listeners to sleep, changing colour with each song, then defecates - and anyone incautious enough to be underneath is turned to stone.

Which leads to the Google Translate weirdness. "Ibong Adarna" means "Adarna bird" ("Adarna" is a proper name of, as far as I can tell, unknown etymology). But if you put it into Google Translate, it correctly detects the language as Filipino (the prestige register of Tagalog), but translates the whole phrase as "Toilet Slave" [or did so at the time of the screen shot below - it has now been fixed].

This is peculiar, to say the least. With the individual words, it translates "Ibong" as "Birds" (which is on the right track) and can't translate "Adarna" (which is expected). The problem is only with the phrase "Ibong Adarna". Is it a malicious mistranslation someone submitted to the database? Some inexplicably garbled allusion to the defecating Adarna bird? Or what? Knowing zero about the Filipino language and the workings of Google Translate, I can't fathom it. Any thoughts?

- Ray (finder's credits to Pinkie17 on Yahoo! Answers)

Addendum: Mark Liberman at the linguistics weblog Language Log kindly mirrored the topic - Really lost in translation - and the Google translation has now been fixed. The mystery still remains as to where the original translation came from.

Tuesday, 21 May 2013

Gormless protoplasm

I just had a brief e-mail exchange with a correspondent (US, I think) who was very amused by my use of the word "gormless", not having encountered it before (no reason to - it is a Britishism). A word I've known from childhood, I love it for its sheer insultingness - its implication not merely of stupidity, but of inept and pitiably-baffled stupidity - and the discussion prompted me to look into the background.

The Oxford English Dictionary tracks it back to the dialect form gaumless / gawm(b)less, where the "gaum" part means "heed, notice, understanding" and goes back to Old Norse gaum-r (masculine), gaum (feminine). From that, I assumed "gormless" must be an old word on the decline, but Google Books Ngram Viewer produced a surprise ...


click to enlarge - gormless 1870-2000
... which is that "gormless" has seen a steady rise in print use since around 1920, taking off dramatically from the competing dialect forms "gawmless" and "gaumless":

click to enlarge - gormless,gaumless,gawmless 1870-2000
Why the word should suddenly catch on in the mainstream is anyone's guess (see the addenda, below, on this). Why it should catch on in that spelling is perhaps more explicable, in that "gaum-" and "gawm-" are pretty peculiar in terms of standard English orthography, whereas "gorm-" (pronounced /ɡɔːm/ in the non-rhotic RP English) is  a normal-looking way of rendering the sound very well. With "gaum" coming from Old Norse, you'd expect "gormless" and its precursors to have a Northern English origin - a relic of the Danelaw era - and they do. Quite apart from appearing in Wuthering Heights ...
Did I ever look so stupid: so "gaumless," as Joseph calls it?
- Heathcliff
... the "gaumless" / "gawmless" forms appear from the early 19th century in a number of northern English regional dialect glossaries and the occasional work of regionally-set fiction (Google search on "gaumless" OR "gawmless").

I thought for a moment I'd beaten the OED's first citation (1883) for the form "gormless" - Google Books produced a handful of earlier hits. But I soon found that the majority of 19th century hits arose from a murk of amusing optical character recognition errors, mostly for "germless":
But there are real occurrences only a little later than the OED's, such as the English Dialect Society's 1886 documentation of "GORMLESS, adj. dull, stupid" in Stockport dialect (see A Glossary of Words Used in the County of Chester, Robert Holland, Pub. for the English Dialect Society, by Trübner & Co., 1886, Internet Archive ID aglossarywordsu00hollgoog).

Addendum:
By complete coincidence, this ties in with the current Language Log post by Mark Liberman, Ngram morality. When Google Ngram Viewer was launched - see Google Books N-gram - wow! - one of the aspects that was hyped was its potential for "culturomics": quantitative research into social trends as reflected in language: this was outlined in the paper Quantitative Analysis of Culture Using Millions of Digitized Books (Science, 14 January 2011: Vol. 331 no. 6014 pp. 176-182).

This is a powerful idea, when the results are interpreted with a lot of caution (I described previously - When pufh comef to fhove - how an apparently robust pre-Victorian era when it was OK to use the word "fuck" in print is entirely an artifact of the word "suck" being printed with a long-s as "ſuck").

But as Professor Liberman and others have discussed, there are those, often seeking confirmation for some world-view, who are ready to wade in with no such caution. One of the dubious forms of analysis is the completely simplistic conclusion that the frequency of a concept mentioned in print is a direct indicator of how that concept applies in society. By that line of argument, the steady rise of "gormless" over the past century means society has become more gormless over that time.

Ngram morality looks at a current op-ed column by the NY Times pundit David Brooks, who applies precisely the same reasoning, based on several like-minded papers, to conclude that society is going to the dogs, as evidenced by the rise and fall of certain words.

Addendum 2:
Martyn Cornell of Zythophile has offered in the comments a theory on the rise of "gormless".
It may or may not be a coincidence that the rise of "gormless" begins at about the same time as the rise of BBC radio: could it be because Northern English comedians were introducing the word to southerners, who took it up with enthusiasm? More research needed ...
This looks a very good start. I don't have any evidence of his using it, but the comedy persona of the immensely popular George Formby was regularly described as "gormless".

- Ray

Thursday, 26 January 2012

When pufh comef to fhove

I just saw an interesting example of the kind of analysis that needs caution when using Google Books Ngram Viewer.

On Yahoo! Answers there was a question asking what word people used instead of "push" before 1800.  This slightly odd query came on the evidence of the Ngram Viewer graph, which appears to show virtually no use of the word before the very late 1700s, then a sudden rise into significant use around 1800.

"push" 1750-2008

The explanation is simple, though the precise timing is hard to explain. Prior to 1800 or so, printed texts used the "long s" character "ſ" (a.k.a. "medial s" or "descending s"), which Google's OCR algorithm interprets as "f". So if you look instead for "pufh", you find all those missing pre-1800s examples of "push". The transition between the two is striking ...

"push" / "pufh" 1750-2008
"push" / "pufh" 1750-2008 (detail)
... and I'm not sure if anyone knows what was happening in the publishing/printing world to account for such a rapid shift.  As the Wikipedia article describes, it happened at different times in different countries, but just as rapidly as in English.

This phenomenon also explains the strange bipolar Ngram Viewer graph for "fuck".

"fuck" 1750-1830

The post-1960 hits are real. The pre-1800s ones don't represent some robust pre-prudish age, but occurrences of "suck" printed as"ſuck".

- Ray

Monday, 17 October 2011

Brigham Young corpus tools

I've a lot of use of Google Books Ngram Viewer since it was launched: an excellent system for searching the whole Google Books corpus (for example, for finding when a word first came into the language). It has some limitations: occasionally poor metadata and OCR errors; the inability to select context (e.g. a search comparing the verb forms "smelled" and "smelt" will be contaminated by the fish called "smelt"); the purely graphical interface; the not-very-useful automated choice of time slots for the search links it generates; and the failure to show words at all if the frequency is too low.

Many of these problems have been overcome in the extremely nice unofficial interface to the Google corpus designed by Mark Davies, Professor or Corpus Linguistics at Brigham Young University: the Google Book American English Corpus ("155 billion words, 1810-2009"). It's very powerful and versatile:

This improves greatly on the standard n-grams interface from Google Books. It allows users to actually use the frequency data (rather than just see it in a picture) ...
...
This interface allows you to search the Google Books data in many ways that are much more advanced than what is possible with the simple Google Books interface. You can search by word, phrase, substring, lemma, part of speech, synonyms, and collocates (nearby words). You can copy the data to other applications for further analysis, which you can't do with the regular Google Books interface. And you can quickly and easily compare the data in two different sections of the corpus (for example, adjectives describing women or art or music in the 1960s-2000s vs the 1870s-1910s).

It's disappointing that it's currently limited to US English, but this is early days.

Note however that what you see here is just a very early version of the corpus (interface), and many features will be added and corrections will be made over the coming months. Also, in June 2011 we applied for a grant to integrate other Google Books collections into our interface, including British English, English texts from the 1500s-1700s, and texts from Spanish, German, and French. If funded (we'll receive word on this in December 2011), each of these additional corpora will be at least 50 billion words in size.

The other corpora by Professor Davies are, despite being considerably smaller, nevertheless still worth exploring. They include: Corpus of Contemporary American English (COCA), Corpus of Historical American English (COHA), TIME Magazine Corpus of American English, BYU-BNC: British National Corpus, Corpus del Español, and Corpus do Português. See the main entry page: CORPUS.BYE.EDU.

- Ray (via Language Log)

Saturday, 18 December 2010

Google Books N-gram - wow!

Google just blew my bibliographic socks off.

Geoff Nunberg at Language Log (see Humanities research with the Google Books corpus) just posted news and some links concerning the Books N-gram Viewer that just went live.

I've enthused previously about the power of Google Books to hack into historical texts in a way that would have been impossible less than a decade ago. The Books Ngram Viewer adds to this facility with a powerful search interface that accesses a humungous corpus of texts (the English one, for instance, covers 360 billion words) and can graph, singly or in comparison, normalised frequencies. The possibilities are immense. As the Science research article abstract says:
We constructed a corpus of digitized texts containing about 4% of all books ever printed. Analysis of this corpus enables us to investigate cultural trends quantitatively. We survey the vast terrain of "culturomics", focusing on linguistic and cultural phenomena that were reflected in the English language between 1800 and 2000. We show how this approach can provide insights about fields as diverse as lexicography, the evolution of grammar, collective memory, the adoption of technology, the pursuit of fame, censorship, and historical epidemiology. "Culturomics" extends the boundaries of rigorous quantitative inquiry to a wide array of new phenomena spanning the social sciences and the humanities.
- Quantitative Analysis of Culture Using Millions of Digitized Books, * Michel, et al. Science 1199644DOI:10.1126/science.1199644
What this actually means, even at a trivial level, is that anyone can do linguistic studies that would have taken years (or even be impossible). You can chart the continuous decline of "whom" in British English over nearly two centuries.  You can compare change of acceptable usages with time: for instance, Mohammedan vs Moslem vs Muslim or Esquimaux vs Eskimos vs Inuit.  You can get long-term statistics for inflected vs. periphrastic versions of adjective comparisons: e.g. "pleasanter" vs. "more pleasant". You can plot the fate of variant spellings: such as how "focused", originally a minority spelling, overtook "focussed" around 1900 and came to dominate it.  You can check age of words: for instance, how people have been "holidaying" (often taken to be a neologism) since 1840. You can look at the history of coexisting forms, such as "none of us is" vs "none of use are" or "Devonshire" vs "Devon".  This is a delight for lexicographical enthusiasts.

The setup isn't perfect. As I and others have mentioned, bad metadata and OCR errors can be a problem. For example, Mark Liberman's follow-up at Language Log - More on "culturomics" - mentions how attempts to trace the history of the word "fuck" in print (see graph) are confused by the "long s", so that pre-1820 you're actually finding occurrences of the word "suck" (like this). More fundamentally, though, once you get away from raw lexical observation and into sociological analysis - the "culturomics" part - it shouldn't be forgotten that frequency of appearance in books is a merely a proxy for the multiple social factors driving that frequency. It would be, for instance, an unreliable conclusion that the British have steadily become less interested in love over the past two centuries because the word's appearance in print has more or less continuously declined.

Nevertheless, searches I've tried often reveal striking patterns, even if they may be inexplicable. Why the seemingly cyclic book references to red sunsets? Why have references to Sherlock Holmes steadily grown over the 20th century? Why do occurrences of the word "fat" rise steadily from 1840 to peak in the late 1870s? What do the peaks in references to opium mean? (this one can be partially answered; two of them coincide with the Opium Wars).  Why are there two 19th century peaks for "Batman" (it seems to be a confluence of coverage of people withthat surname, notably John Batman). Does the post-1960 rise in references to "Frankenstein" mean anything culturally or does it just reflect the success of particular movies. I have a feeling I'm going to be making a lot of use of this.

The Guardian has a more general piece on it here: Culturomics and the new Google tool for tracking cultural trends. See also the official Culturomics site.

Addendum: the paper Quantitative Analysis of Culture Using Millions of Digitized Books (Science DOI: 10.1126/science.1199644) - free registration with Science is required - is very worth reading. It mentions some highly interesting areas including:
  • The recent massive growth in the English lexicon (over 70% during the last 50 years).
  • The trade-off of dictionaries in balancing comprehensiveness and conciseness, with the result that over half of the English lexicon comprises "dark matter" that doesn't appear in dictionaries.
  • The ability to track trends such as the regularisation of verbs, such as the shift from "-nt" endings to "-ned" (e.g. "burnt" to "burned").
  • The characteristic trajectories of appearance in print as a proxy of fame.
  • Detection of censorship by non-appearance in print: notably the absence from German texts of individuals identified as undesirables under the Nazi regime.
  • Culturomics - the identification of "fossils" of cultural trends through print frequency (e.g. "influenza" being mentioned a lot in print at the time of known pandemics).

The epidemiology example illustrates an important limitation to "culturomics". As quoted in Wired:

Patterns that can be queried from its cloud are not necessarily answers unto themselves, they say, but a way of illuminating subjects for further investigation.

"It’s not just an answer machine. It’s a question machine," said study co-author Erez Lieberman-Aiden, a computational biologist at Harvard University. "Think of this as a hypothesis-generating machine."
- Cultural Evolution Could Be Studied in Google Books Database, Wired, Dec 16th 2010

A look at the references to "cholera" shows peaks that may correspond to epidemics, but the largest, in the mid-1880s, more likely corresponds to the topicality of Robert Koch's isolation of Vibrio cholerae in 1884.

Addendum: discussion at Language Log - see True Grit isn't true - highlighted another significant problem with the setup.  For some reason (maybe to do with OCR, indexing, tokenization or the search interface) Google Books N-gram Viewer seemed to underestimate by three to four orders of magnitude (!) occurrences of forms with apostrophes. This made it useless for examining historical occurrences of contractions in English. Correction: see Google n-gram apostrophe problem fixed.

Ray

Saturday, 10 July 2010

William Barnes online (depending where you are)

Onwith from the posts yesterday and backalong - as I guess William Barnes might like me to write (see Ansible ... and Anglish), I've just been reading his An Outline of English Speech-craft (CK Paul, 1878).

Sunday, 14 March 2010

Street View


View Larger Map

Ooh, look: the bookshop where I work is on Google Maps Street View, which has just been updated to include Exeter. You can take a tour around Topsham if you like. I don't recall seeing the Google car, but looking at the orange window display with the Halloween pumpkins, the photo would have been around the end of October.

- Ray

Thursday, 11 March 2010

Believe you me

Nice example of the power of Google Books as a tool for conducting fairly novel linguistic research. I just ran into a question on Yahoo! Answers about the origin of the expression "Believe you me", with its unusual verb-subject-object (VSO) construction.

Michael Quinion at the excellent World Wide Words shows evidence - see Believe you me - that English has an archaic VSO form used for imperatives, as in the King James Version of the Bible

For thus saith the Lord unto the house of Israel, Seek ye me, and ye shall live

but mentions that "Believe you me" is an outlier that came into the language late.

For further help here, I turned to Benjamin Zimmer, at the University of Pennsylvania, an ace at researching historical word usage. He tells me that there are earlier examples, but that nearly all of them are in verse, where the phrasing is useful for scansion. He has been able to find only three examples in prose from the nineteenth century.

What seems to have happened is that a once-standard phrase that had been lurking in the language for generations suddenly became much more popular and widespread around the 1920s. What we have here is a revitalised fossil, a semi-invented anachronism

It's amazing how access to sources has blossomed since that was written in 2006. Google Books now finds dozens of prose hits from the 19th century:

  • "Well, then, believe you me, that all poor Biddy's hope rests on her sweet Saviour" - 1829
  • "Master Denis shall feel, believe you me, what it is..." - 1829
  • "and believe you me, before that day twelvemonth it. would take more nor a yard of tape to make them an apron-string" - 1834
  • "it involves misery for themselves and for their children, believe you me the same selfish instincts that prompt these men..." - 1839
  • "The work is God's, and not man's; and believe you me, that not only Annas and Caiaphas, but Herod and Pilate too, are against Luther" - 1840
  • "Well, sir, believe you me, I'll give that lassy as good a strapping as ever she got when she comes back" - 1841
  • "for, believe you me, I'll never look at the same side of the road with Tade Ferrall again" - 1842.
  • "But believe you me, they will not be able to proceed much further" - 1847
  • "No, no, believe you me, sir, it won't do in a free country" - 1857
  • "Helter-skelter, pell-mell down the staircase they flew, and, believe you me, the commodore helped them in their descent" - 1861
  • "Believe you me, that if you were blown down there you'd be bruised to mummy among the tombstones" - 1864
  • "But many a time, believe you me, I bought a rabbit from you" - 1870
  • "Balzac stands for Paris, believe you me" - 1877
  • "We've not come to the worst yet, believe you me" - 1877
  • "Believe you me, they wouldn't have done the like of that without they'd got their orders direct from the Castle" - 1878
  • "No, Mike, believe you me, while grass grows or water runs" - 1882
  • etc ...

The interesting observation is, at first glance, that the majority of these early examples come from Irish publications: The Christian examiner and Church of Ireland magazine, Sketches in Erris and Tyrawly, The Dublin University Magazine, the Irish novel In Re Garland: A Tale of a Transition Time, Michael Banim's The Town of the Cascades, and so on. This could explain why "believe you me" is an outlier by date. A hypothesis: perhaps it isn't a relic of the archaic English VSO construct, but arrived by a different route. Irish Gaelic is a VSO language; maybe its sentence pattern influenced Irish English? A quick Google suggests "Believe you me" might correspond to the Irish phrase "Creid uaim é!" = "Take it from me!" (literally, "believe from-me it").

The 1857 example, by the way, isn't Irish, but Charles Dickens editorial in Household Words magazine.
- Ray

Monday, 15 February 2010

Snowclones are born, not made

I've enthused before about the power of the Internet, Google Books in particular, at making available a vast corpus of text that enables linguistic searches to be conducted in a moment that would have taken decades of research in pre-Internet days. It's possible, for instance, to rapidly verify dates for the existence of a word: very handy when you run into "recency illusion", the term coined by linguist Arnold Zwicky for the mistaken belief that words or expressions are new. The easy searchability of online texts has given rise to the creation and study of whole new classes of linguistic observations, such as the "snowclone", a term coined by Glen Whitman on January 15, 2004 and rapidly adopted in linguistics circles. The snowclone is a boilerplate cliché, alluding to a classic example "If Eskimos have N words for snow, X surely have Y words for Z", consisting of a template that appears repeatedly with various different words slotted in. See The Snowclones Database for many examples.

I spotted a snowclone yesterday in relation to an enquiry on Yahoo! Answers about the origin of the expression "Leaders are born, not made". It turned out the expression had no clear origin, and appeared instead to be one variant of a snowclone that dates back at least to the 1700s.

The story seems to have started with the aphorism "Poets ... are born, not made ...". This example is by Aphra Behn, 1714, but she's only one citer of this particular form of the aphorism expressed in Latin as Poeta nascitur non fit. Despite being frequently attributed to Horace or Cicero, it doesn't appear in any classical text. See Poeta Nascitur Non Fit: Some Notes on the History of an Aphorism, Journal of the History of Ideas, Vol. 2, No. 4 (Oct., 1941). This was the prototype and the predominant version in the 1700s; but then it snowballs. The snowclone started out as "Xs, like poets, are born, not made" and then the comparison to poets was gradually dropped in favour of "Xs are born, not made". Let's see some historical sightings.
  • "Thou hast a Genius, and a Swinger; Thou'rt Born, not Made a Ballad-singer" The Athenian Oracle, 1709
  • "Actors, like poets, must be born, not made" Mrs Griffiths, 1775
  • "... the historian, like the poet and the orator, must be born, not made" Tobias Smollett, 1783
  • "Like poets, historians are born, not made." Francois Xavier Martin, 1827
  • "Religions are born, not made" Alonzo Hill, 1831
  • "what is said of the poet is also thought of the philosopher — that he is born, not made" Willian Greer, 1832
  • "painters .. are born, not made" Unknown, The American Monthly Magazine, 1834
  • "rope-dancers — they are born, not made" Henry Junius Nott, 1834
  • "Still that which is accounted true of poets holds equally good of pickpockets—who are born, not made;" White & Meadows, 1838 1
  • "Genius: born not made" Thomas W Dorr, 1841
  • "the true orator is 'born, not made'" Sydney Smith, 1841
  • "Whigs, like Poets, are born, not made" Unknown, 1845
  • "A thoroughly vulgar person is — like the poet — born, not made", Chambers Edinburgh Journal, 1846
  • "The number, however, of those who are capable of discovering scientific principles is comparatively small; like the poet, they are 'born, not made'" Professor Henry, 1848
  • "A gentleman is like a poet — he is born, not made" Thomas Hood, 1848
  • "Colorist, the, is born, not made such" Osborn & Bouvier, 1849
  • "an editor must be 'born, not made'" Anon., 1850
  • "republicans are born, not made" William Starbuck Mayo, 1850
  • "A good housewife, like a good poet, is "'born, not made'" Eliza Cook, 1852
  • "a wit is born, not made" Yale Literary Magazine, 1857
  • "grooms, like poets, are born, not made" George Borrow, 1857
  • "The highest military genius, as the highest poetical genius, is born not made." Dublin University Magazine, 1857
  • "Housekeepers, like editors and poets, are born, not made", 'Hester', 1858
  • "angler must he be born, not made" P.P., 1860
  • "commanders are born, not made." Mr Fessenden, 1861
  • "Teachers are born, not made" Professor Wickersham, 1862
  • "Really good talkers are born, not made" The Continental Monthly, 1862
  • "true kings, like true poets, are born, not made." London Quarterly Review, 1862
  • "To make my position — that botanists must be born, not made — doubly sure" Botanical Society of Edinburgh, 1863
  • "but the true chopper, like the true poet, is 'born, not made.'" Journal of Horticulture and Gardening, 1865
  • "It is a mistake to think good servants, like poets, are born, not made" Cornhill Magazine, 1866
  • "Leaders are born, not made" Freewill Baptist Quarterly, 1867
  • "Now, real educators are born, not made" Liberty magazine, 1867
  • "The architect, as an artist, is born, not made" William Laxton, 1867
  • "mannerly people, like poets, are born, not made," Rhoda Broughton, 1868
  • "Wardens are 'born, not made'" The Methodist Review, 1868
  • "A good cook is born, not made, but he needs an immense deal of polishing" Putnam's Magazine, 1869
  • "Good bread makers, male or female, are born, not made" The Overland Monthly, 1869
  • "a dog-driver, like a poet, is born, not made" George Kennan, 1871
  • "the great mechanic, like the great poet, is born, not made" Samuel Smiles, 1884
  • "The Christian in short, like the poet, is born not made" Henry Drummond, 1887
  • "The genuine secretary is born, not made" Grant Allen, 1889
  • "snobs are born — not made" Mary Elizabeth Wilson Sherwood, 1897
Bored now. Yes, it seems anybody with any kind of talent or distinctiveness is... And that's not exhaustive, nor even using up 19th century examples; a nice example of a snowclone that has persisted for around three centuries.

1. I wonder about the White & Meadows comment on pickpockets. Dashiell Hammett, in his 1923 From the Memoirs of a Private Detective, said of it: "Pocket-picking is the easiest to master of all the criminal trades. Anyone who is not crippled can become adept in a day".

PS: Thanks, kalebeul: cool. Indeed.

- Ray

Out-takes: copyright and crime

I've been getting into a daft habit of carefully making a note of news relevant to JSBlog, then forgetting to upload it. Purged from my organiser:

Men at Work plagiarised 'Down Under' riff ("Flute melody taken from 1935 'Kookaburra' children's song, Australian court rules", Kathy Marks, The Independent, 5 February 2010). This means Men at Work potentially owe millions to the copyright owners, Larrikin Music. See the previous Kookaburra fossil exposed for background. Personally I think the result stinks, and that the quotation in question, a tiny riff between verses, was nothing more than a nice homage to a tune that had become de facto public domain due to its obscure copyright status. Quoted in The Age, the founder of Larrikin Records and original owner of Larrikin Music, Warren Fahey, says exactly this:
He recommends the copyright owner, Larrikin Music, should "gift" the song to Australia, arguing that most Australians believe they already have public domain ownership of the song anyway.

"The past week has seen thousands of emails, letters to the editor, radio commentary and internet forums criticising the judgment," says Fahey, who sold Larrikin Music to Music Sales Corporation in 1988 and whose folk band is called the Larrikins.

"Many of these incorrectly criticise Larrikin Records and myself as the protagonist, asking, 'How could someone so dedicated to Australian music do such a thing?' The Larrikin brand has certainly been tarnished by what many see as opportunistic greed on behalf of Larrikin Music/Music Sales."
A "larrikin" is, incidentally, one who subscribes to Australia's folk tradition of anti-authoritarianism - see Larrikinism - and the opposite of a "wowser".

Harry Potter and the great Google onslaught ("The bunfight over Google's library project only serves to remind us that intellectual property battles are nothing new", Robert McCrum, The Observer, Sunday 14 February 2010): the always interesting McCrum writes about great intellectual property battles in history such as the first recognition of literary "piracy" in the 1660s, arguing in effect that these didn't kill literature in the past, so the disputed Google Library Project may have innovative effects we can't predict. See preview of Adrian John's book Piracy.

Forget 'serious' novels, I've turned to a life of crime ("Murder mysteries, once looked down on, are now fit for the literary elite", Stephanie Merritt, The Observer, 14 February 2010). Merritt argues the ... er ... merit of the crime genre as historical fiction. I thoroughly agree; one of the best historical novels I know is Peter Lovesey's murder mystery Wobble to Death, with its fascinating background in Victorian endurance races and their attendant corruption and strychnine-doping (see Strychnine - a lesser-known past).

This artwork was made by a killer. It is no less valid for that (Deborah Orr, The Independent, 11 April 2009): concerning the difficult issue of what we should feel about art made by murderers. In the literary field, she mentions the well-known case of William Chester Minor, a schizophrenic murderer who made thousands of contributions to the OED from his cell in Broadmoor. The excellent and prolific historical crime novelist Anne Perry, creator of the William Monk series, also springs to mind.

- Ray

Wednesday, 28 October 2009

Rings of Saturn and other Litmaps

Folowing on from the Bruges-la-Morte post, I still haven't completely read WG Sebald's The Rings of Saturn (I came as close as I ever have to swearing in front of a customer a few weeks back when he bought the single copy I hadn't noticed was on the shelves). Meanwhile, there are continuing interesting posts at the blog Vertigo: Collecting & Reading W.G. Sebald.

I'm very much inspired by the possibilities suggested in Mapping Sebald's Literary Landscape, which links to Barbara Hui's "Litmap" - an annotation of Google Maps with the places mentioned during the East Anglian walking tour of the unnamed narrator of The Rings of Saturn.

See her Litmap Presentation Notes for background. The notes mention a similar project, Gutenkarte, that mines Project Gutenbeg texts to create presentations with placenames hyeprlinked to a map. It has rough edges - for instance, it thinks Providence might be a location in

they had a notion that Providence would interfere in favour of him who was in the right.
- The Journal of a Tour to the Hebrides with Samuel Johnson, LL.D., James Boswell

but it seems a very good interface for reading works with a strongly geographical component.
- Ray

Wednesday, 14 October 2009

Toff: cod etymology and duff metadata

My cod etymology alarm rang when reading yesterday's Western Morning News

Mr Bell may be interested to learnt that "toffs" is an abbreviation of toffee-nosed person, a term arising when snuff was the favoured "stimulant" used by the very rich; the ones who could afford the many varieties and strengths of snuff.

As a pinch of snuff was inhaled into nostril, and because dribbling noses are not uncommon in winter, it appeared as if toffee was falling from the nostrils of the partakers of this "drug" - the rich upper class
- Getting up to snuff on why toffs are so called, Lord Clifford of Chudleigh, Letters, Western Morning News, October 13 2009

Picturesque an image though this is, I'm of the opinion that it's complete bilge. As this Guardian review says, this etymology for the terms appears in Ian Kelly's 2006 biography Beau Brummell, The Ultimate Dandy -

... the origin appears to derive from the unsightly brown droplets that dripped from a gentleman's nose after taking snuff - which of course was only taken by the "upper class"

- but I'll believe that when I see a contemporary source. And that's the key point with origin stories: it's not sufficient that they be plausible; you have to find evidence of their formation. As far as I can find, there's no sign of anyone using the terms historically to refer to snuff-taking (if this usage existed, you'd expect to find it in descriptions such as A pinch of snuff, anecdotes of snuff taking, with the moral and physical effects of snuff, by Dean Snift of Brazen-nose, Benson Earle Hill, 1840).

In fact the evidence is that "toff" predates "toffee-nosed", and neither of the terms are very old. The Oxford English Dictionary's first citation for "toff" is 1851, in Mayhew's London Labour and the London Poor (it appears in the account of crossing-sweepers of how they solicit money from customers - here). The most likely etymology according to the OED is from "tuft", a term for upper-crust Oxbridge undergraduates, who distinguished themselves from the hoi polloi by wearing a gold tuft or tassel on their college caps. The pseudonymous 1854 college novel The Adventures of Mr Verdant Green - see Wikipedia - has a footnote

As "Tufts" and "Tuft-hunters" have become "household words," it is perhaps needless to tell any one that the gold tassel is the distinguishing mark of a nobleman.

John Camden Hotten's 1874 Slang dictionary: etymological, historical, and anecdotal explains "tuft-hunter" as being a social hanger-on who seeks the society of the wealthy. Hotten also has a significant entry for those who might think "tuft" to "toff" an unlikely jump; he lists an intermediate form "toft"

Toft, a showy individual, a swell, a person who, in a Yorkshireman's vocabulary, would be termed "uppish"

On to "toffee-nosed". The OED's first citation for this is 1943.

Toffee-nose, another of the expressions chiefly heard amongst the W.A.A.F. This refers to a snob or someone who considers herself ‘superior’. It is very apt since it implies that the nose is kept high to prevent it coming into contact with the mouth.
- Service Slang, John Leslie Hunt, A. G. Pringle, 1943

"Toffee-nosed", then, is "toffy-nosed": having the nose of a toff. The description "the nose is kept high" evidently refers to body language - the stereotypical nose-high posture indicating contempt for the lower orders - rather than any toffee-like nasal exudation.



Of course, etymologies are always open to question, and things may not be so clear-cut. Nevertheless, the chronology at least is readily verifiable using Google Books. I can find no pre 20th century uses of "toffee-nosed" (or "toffy-nosed" or "toffee nose"). Here's the search for the range 1600-1900: nothing, except three false positives from dodgy metadata. I couldn't resist, however, trying to pre-date the OED 1943 citation, which seemed a bit late, and the earliest example I can find is in a 1914 edition of Punch, where a cartoon features the aftermath of a fight between Boy Scots, with the caption:

The Victor (after being admonished for un-scoutlike behaviour). "Well, you may say what you like, Sir, but I consider it distinctly subversive of discipline for an ordinary private to call his patrol-leader 'Toffee-nose.'
- Punch or the London Charivari, Vol. 147, December 2, 1914. Project Gutenberg.

A 1921 Notes & Queries lists "toffee-nosed" as trenches slang equivalent to "stuck up". Nevertheless, it chiefly kicked off during World War II, evidently still rooted in military slang:

A premature 'life' will do more to disgust the select and superior people (the RAF call them the 'toffee-nosed') than anything.
- The Letters of TE Lawrence, Thomas Edward Lawrence, ed. David Garnett, 1939

They're county people, all frightfully toffee-nosed and Poona.
- Pastoral, Nevil Shute, 1944

"You wouldn't know a gentleman if you was ter see one, you toffy-nosed coot!"
- Tinned soldier: a personal record, 1919-1926, Alec Dixon

For various reasons, tracking "toff" via Google Books gets into a mess of metadata and indexing problems. Firstly, a search gives many false positives on the German word "Stoff" in Fraktur typeface, so a good start is to limit the search to English texts. Secondly, as you go into older and older texts, you find results increasingly contaminated by mishits on words vaguely resembling "toff" or "toffs" (such as love, Topp, taff, feoff, Tait's, and so on), making it very difficult to search for occurrences in the pre-1851 slot of interest. I don't know what this means about the digitising/indexing process; it happens with the Times Digital Archive and the British Library Nineteenth Century Newspapers database too. Mishits for "toffs", a species of fragrant thistle, are another sidetrack. Once you get to the late 1800s - see search results for 1870-1900 - the false hits clear up, and the results at least show "toff" appears several decades before "toffee-nosed".

So, no luck with beating the OED on "toff" citations. Still, I found a spectacular Google Books metadata crash en route: this hit whose metadata is for the 1845 Archäologische Aufsätze by Otto Jahn, but whose text is Theodore Watts-Dunton's 1898 novel Aylwin. I also found a lovely Melbourne Punch article for September 3rd 1868, Toffs, which leads with spoof etymologies into a taxonomy of varieties of toff as exotic creatures.


- Ray

Monday, 21 September 2009

The hand of Google

Slight scan failure in Google Books, from The marvellous and incredible adventures of Charles Thunderbolt, in the moon, Charles Rumball, 1851, between pages 189 and 190. (I found it by accident while Googling for literary cheesemites thus).



Coo: I feel like the brother of Jared. I'd never run into it before, but it happens occasionally: see Google Books adds hand scans at TechCrunch. The pink finger cots appear to be standard kit for the job.
- Ray