Friday, April 24, 2009

Healed by the right words

We all know that placebos can be surprisingly effective. But - though it's not exactly surprising - I hadn't realised that there is experimental evidence that simply saying the right thing can have a curative effect.

Two hundred patients with abnormal symptoms, but no signs of any concrete medical diagnosis, were divided randomly into two groups. The patients in one group were told "I cannot be certain what is the matter with you", and two weeks later only 39% were better"; the other group were given a firm diagnosis, with no messing about, and confidently told they would be better within a few weeks. 64% of that group got better in two weeks." (Bad Science, p. 75, citing Thomas 1987)

I can imagine a lot of factors that could affect the effectiveness of the doctor's words here - mainly anthropological, but some of them would certainly fall within the domain of linguistics. For example, the intonation pattern will affect the patient's perception of the doctor's confidence; does that affect the efficacy? Likewise, the accent and the choice of vocabulary could both affect comprehension and perceived competence, and hence presumably the efficacy. Not really my field, but it could be a line of research with unusually clear-cut potential benefits. The obvious problem with this example is that it involves doctors lying to patients, but if the effect could be reproduced without that it would certainly be worth doing.

Bibliography:
Thomas KB. General practice consultations: is there any point in being positive? BMJ (Clin Res ed) (9 May 1987); 294 (6581): 1200-2.

Thursday, April 23, 2009

"Political complexity predicts the spread of ethnolinguistic groups"

An interesting paper: Political complexity predicts the spread of ethnolinguistic groups. Two basically unsurprising claims that it's good to have calculations supporting: "pastoralists were found to have larger language areas than agriculturalists" and "languages associated with more politically complex societies cover significantly larger areas than those of less complex societies". They also present arguments that "although regions of high biological and cultural diversity do overlap to a striking degree, it is unlikely that biological diversity has any direct effect on cultural diversity on a global scale." Surprisingly, mountainousness was found to correlate with larger language areas, not smaller ones - seems a little suspicious that, though some mountainous areas are pretty un-diverse. Flaws: well, it relies on Ethnologue data and GMI maps, both of which are often unreliable, and systematically more splittist in some areas than in others; but it's not obvious that that would substantially affect the result. Also, ethnic groups, languages, and political units very often don't match up, and their measure of political complexity is based on data for ethnic groups rather than for languages.

(Via GNXP.)

Friday, April 17, 2009

A Fulani village in Algeria

Anyone acquainted with West African history will be aware of the remarkable extent of the Fulani diaspora, stretching from their original homeland in Senegal all the way to Sudan. However, I was surprised to read the following note in a history of the Tidikelt region of southern Algeria (around In-Salah):
"Le village actuel de Sahel a été créé en 1779 par Sidi Abd el Malek des Foullanes, venu à Akabli dans l'intention de se joindre à une pèlerinage, dont le départ n'eut pas lieu... Les Foullanes sont des Arabes originaires du Macena (Soudan); il y a encore des Foullanes au Sokoto; Si Hamza, le cadi d'Akabli appartient à cette tribu." (L. Voinot, Le Tidikelt, Oran:Fouque 1909, p. 63)
(The current village of Sahel was created in 1779 by Sidi Abd el Malek of the Fulani, who had come to Akabli with the intention of joining a pilgrimage whose departure never occurred... The Fulani are Arabs originating from Macina (Sudan [modern-day Mali]); there are still Fulani at Sokoto; Si Hamza, the qaid of Akabli, belongs to this tribe.)

I very much doubt there would be any traces of the language left - even assuming that Sidi Abd el Malek came with a large enough entourage to make a difference - but wouldn't it be interesting to check?

Sunday, April 12, 2009

How many words are there in a language?

In a recent discussion, the question came up of whether a language's vocabulary could be tallied (briefly addressed at Language Log a while back, and at FEL.) I have no firm answer to that (and it's logically independent of whether or not you can estimate the proportion of the vocabulary coming from a given language - that's a sampling problem.) But, notwithstanding the bizarre if occasionally entertaining acrimony of that discussion, it's actually a rather interesting question.

Clearly, any given speaker of a language - and hence any finite set of speakers - can know only a finite number of morphemes, even if you include proper names, nonce borrowings, etc. ("Words" is a different matter - if you choose to define compounds as words, some languages in principle have productive systems defining potentially infinitely many words. The technical vocabulary of chemists in English is one such case, if I recall rightly.) Equally clearly, it's practically impossible to be sure that you've enumerated all the morphemes known by even a single speaker, let alone a whole community; even if you trust (say) the OED to have done that for some subset of English speakers (which you probably shouldn't), you're certainly not likely to find any dictionary that comprehensive for most languages. Does that mean you can't count them?

Not necessarily. You don't always have to enumerate things to estimate how many of them there are, any more than a biologist has to count every single earthworm to come up with an earthworm population estimate. Here's one quick and dirty method off the top of my head (obviously indebted to Mandelbrot's discussion of coastline measurement):
  • Get a nice big corpus representative of the speech community in question. ("Representative" is a difficult problem right there, but let's assume for the sake of argument that it can be done.)
  • Find the lexicon size required to account for the 1st page, then the first 2 pages, then the first 3, and so on.
  • Graph the lexicon size for the first n pages against n.
  • Find a model that fits the observed distribution.
  • See what the limit as n tends to infinity of the lexicon size, if any, would be according to this model.


A bit of Googling reveals that this rather simplistic idea is not original. On p. 20 of An Introduction to Lexical Statistics, you can see just such a graph. An article behind a pay wall (Fan 2006) has an abstract indicating that for large enough corpora you get a power law.

But if it's a power law, then (since the power obviously has to be positive) that would predict no limit as n tends to infinity. How can that be, if, for the reasons discussed above, the lexicon of any finite group of speakers must be finite? My first reaction was that that would mean the model must be inapplicable for sufficiently large corpus sizes. But actually, it doesn't imply that necessarily: any finite group of speakers can also only generate a finite corpus. If the lexicon size tends to infinity as the corpus size does, then that just means your model predicts that, if they could talk for infinitely long, your speaker community would eventually make up infinitely many new morphemes - which might in some sense be a true counterfactual, but wouldn't help you estimate what the speakers actually know at any given time. In that case, we're back to the drawing board: you could substitute in a corpus size corresponding to the estimated number of morphemes that all speakers in a given generation would use in their lifetimes, but you're not going to be able to estimate that with much precision.

The main application for a lexicon size estimate - let's face it - is for language chauvinists to be able to boast about how "ours is bigger than yours". Does this result dash their hopes? Not necessarily! If the vocabulary growth curve for Language A turns out to increase faster with corpus size than the vocabulary growth curve for Language B, then for any large enough comparable pair of samples, the Language A sample will normally have a bigger vocabulary than the Language B one, and speakers of Language A can assuage their insecurities with the knowledge that, in this sense, Language A's vocabulary is larger than Language B's, even if no finite estimate is available for either of them. Of course, the number of morphemes in a language says nothing about its expressive power anyway - a language with a separate morpheme for "not to know", like ancient Egyptian, has a morpheme for which English has no equivalent morpheme, but that doesn't let it express anything English can't - but that's a separate issue.

OK, that's enough musing for tonight. Over to you, if you like this sort of thing.

Houhou yentakheb rouhou


(Warning: this post contains no significant linguistic content.)

The results are in: Bouteflika has been “re-elected” as President of Algeria with a staggering 90.24% of votes cast. According to Government figures, 74.54% of eligible voters voted (although oddly enough, the polling booths looked deserted in all the main towns.) He had already served two terms, which had been the limit, so, to let himself run for re-election, he had had the constitution changed shortly beforehand. I would start mocking the guy, but why bother? With figures like that, he's making a fool of himself with no help from me. Time was when he was willing to settle for figures that naive observers might be capable of taking seriously; as he turns senile either his intelligence or his capacity for shame must be declining. The best measure of the glory of his achievements is the 50% of Algerian youths who intend to try to leave the country.

In case you were wondering how this result was achieved, here's my best somewhat informed guess: In the countryside, especially in areas like the Sahara where tribalism is still present, the local patriarchs simply tell everyone to vote en masse for the President, on the basis that he will stay in power no matter what they do and a conspicuous display of loyalty will earn them government investment (although even that wouldn't be enough to produce things like the 97% turnout in Tissemsilt without further fraud.) In the cities or the larger towns of the north, practically nobody bothers to vote apart from people on government payrolls, so they simply exaggerate the participation figures. In Kabylie, uniquely, we have a largely rural, somewhat tribal region fed up enough with the government that even the villages have organised themselves to refuse it legitimacy, so conspicuously that even government figures acknowledge a much lower turnout. If we assume that the government figures are broadly accurate regarding relative turnout (though certainly not absolute), then the situation shows up in the negative slope on this plot of population against turnout (participation); the two 30% wilayas are Tizi-Ouzou and Bejaia, the main Kabyle regions.



Another post on this worth looking at: Victory over the People.

Wednesday, April 08, 2009

When goals create blind spots

You're watching a ball game attentively. A person in a gorilla suit walks right through the middle, remaining visible for 5 seconds. Can you imagine not noticing the gorilla guy? Well, it turns out that nearly half of all people undergraduate volunteers don't, if they're busy trying to count passes - and the authors of that study cite 7 other experiments confirming the same principle.

It strikes me that there's a lesson there for linguists. Often linguists study a language for a specific theoretical goal - looking at Malagasy primarily to see what VOS syntax is like, or Oneida primarily to learn how polysynthesis works, or Songhay primarily to see whether it's related to Nilo-Saharan or not. That's fair enough; no one can focus on everything at once. But we can miss some really interesting stuff by focusing on one aspect of the language to the exclusion of others. For example, when Laoust studied Siwi, he was interested almost exclusively in its Berber origins - and as a result, his generally excellent study somehow ignored the vowels e and o (which are found even in Berber words, but are not phonemic in the Moroccan Berber varieties he was more familiar with), and mistakenly attributed the Arabic elements of Siwi to the adjacent Bedouin dialects, when in fact they show some very distinctive non-Bedouin characteristics. This is something we all need to watch out for.

Sunday, April 05, 2009

Flora of the Central Sahara and elsewhere

Ever found yourself trying to sort out a plant name you've elicited, not knowing any botany worth mentioning? Well, it turns out the botanists are a step ahead of the linguists on the digital libraries game, at least in Spain: the Digital Library del Real Jardín Botánico CSIC has a pretty remarkable array of books to browse online. The one that just saved my etymology of the Kwarandzyey plant name tsifəṛfəẓ is Etudes sur la flore et la végétation de la Sahara centrale. Vol. III: Hoggar, which gives both Tamasheq and binomial names for each plant mentioned. Unfortunately it's clear that not all the works give translations of the names, but it's still worth a look.

On a similar note, I've found Sahara-Nature handy sometimes.

Thursday, March 19, 2009

Beni-Snous: Two unrelated phonetic forms for every noun?

I got flabberghasted recently by a casual statement in Destaing (1907:212)'s grammar of the Berber dialect of Beni Snous in western Algeria (near Tlemcen). I nearly missed it as I skimmed it; see if you can spot it. (The translation is mine, as are the bits in brackets.) All the numerals above 1 are from Arabic here, but that's nothing surprising - the same is true in Tarifit, and few Berber varieties have retained the numbers above 3.
"The numbers from 2 to 9 inclusive are followed by the Berber noun in the plural [eg]:

two men ..... θnāịẹ́n ịírgǟzĕn
six women ... sttá n tsénnạ̄n
[...]

From "10" to "19" inclusive, the number is followed by the Arabic singular substantive:

eleven women ... aḥdăɛâš ĕrmra (Algerian Arabic mṛa "woman" مرة; contrast Beni Snous Berber θä́mĕṭṭūθ "woman")
fifteen cows ... ḫamstaɛâš ĕrbégra (Algerian Arabic bəgṛa "cow" بڨرة)
sixteen mares ... sttɛâš ĕrɛấuda (Algerian Arabic `əwda "mare" عودة; contrast Beni Snous Berber θáimārθ "mare")

After the number nouns "twenty, thirty, forty" etc., one uses the Arabic substantive[...]

twenty women ... ɛašrîn ĕmra
fifty mules ... ḫamsîn beγla (Algerian Arabic bəγla بغلة "mule")

a thousand rams: âlĕf kebš (Algerian Arabic kəbš كبش "ram"; contrast Beni Snous Berber išérri "ram")"

If I thought it were remotely possible for Destaing's claim to be true of counting every noun in the language - rather than, say, just the six nouns he gives appropriate examples for - I would be putting together an application to head out to Tlemcen instead of making this posting. (I might still do that anyway some time, mind you.) But for rather a lot of minority languages, all or nearly all speakers are bilingual. And if all speakers are bilingual, what in principle is there to prevent the grammar from containing a rule like this?

So I ask: have you ever come across anything similar elsewhere?

Wednesday, March 18, 2009

Scanned Multi-Alphabet Arabic Manuscript Online

The Princeton Digital Library of Islamic Manuscripts has put a large number of Arabic, Persian, and Turkish scanned manuscripts online. Plenty of interesting stuff there, but one that particularly stood out for me was the untitled Treatise on ancient, alchemical and magical alphabets. Behold the Omniglot of its day! (Well, it's apparently only from the 1700s, but probably a copy of an older work.) It gives tables for the supposed alphabets of each prophet, with the letter names on one page and the letter forms on the next. I'll just point you to a few of the highlights:

Knowing my readers, I suspect I'll have identifications of several of the alphabets I didn't recognise coming soon - although many, perhaps most, of them are certainly made up. Extra points for anyone who can come up with a picture of a magic bowl or something actually using one of the made-up alphabets.

Two other Arabic manuscripts there of potential interest: The conquest of Africa, from Qayrawan to Zab; Book of the Roman months.

Wednesday, March 11, 2009

išni: a Berber ovine, or a Songhay goat?

In Kwarandzyey (Tabelbala), the non-specific word for a sheep or goat is išni. It looks kind of Berber, and the words for different ages or sexes of sheep and goat are definitely from Berber, so I had assumed it must be Berber. But I've never found a term like it in any Berber dictionary. Maybe some reader will tell me that the word is familiar from his/her own hometown, but I just realised that there's an alternative explanation...

The word for "(female) goat" across Songhay may be reconstructed as *hìnčìnì (Nicolai 1981 gives *hìnkìnì, but in all the Songhay languages he cites except Kwarandzyey, original *k and *č both turn into the same sound before front vowels.) Nicolai 1981 gives amkkən "male goat" as the Kwarandzyey reflex of this word, but in fact (as Kossmann first pointed out to me) that turns out to be another one of the Berber etymologies that only Zenaga seems to explain: ämkän "jeune bête (tout animal de pâturage)" (Taine-Cheikh 2008). Instead, I'd like to propose that išni is the Kwarandzyey reflex.

*n is occasionally lost in Kwarandzyey (eg gwa "see" < *guna); I don't know any rule for this so far, but here it might be motivated by dissimilation. Initial *h is lost fairly commonly (at least "water", "man", "two", "three", "hunger"), so that's not necessarily a problem. Short vowels, most commonly (but not always) *i and *u, are frequently deleted, according to a rule whose conditioning I've been investigating lately. *č regularly becomes ts, but when immediately followed by a consonant regularly simplifies to s for all but some of the most conservative speakers. And s and š are not phonologically distinct (except for younger speakers, under heavy Arabic influence); the consistent use of š here would be explained by the i's flanking it. So that would yield *hìnčìnì > *inčni > *itsni > isni = išni.

Of course, if išni is attested in Berber then all this reasoning may have to be rethought - so if you speak Berber and have heard the word before, please tell me now!

Arabic (and Berber?) loanwords in southern Italy

Just came across a little monograph on Arabic and Berber loanwords in the dialects of the Basilicata (southern Italy): Sopravvivenze lessicali arabe e berbere in un'area dell'Italia meridionale, la Basilicata by Luigi Serra. Most of the loans listed are from Arabic, some quite obvious (eg taūt "coffin" < تابوت, źir "a copper or terracotta container for liquids" < زير, zammîl "big pannier with which various goods are transported on a beast of burden's back" < زنبيل), others rather less clear-cut.

Only three loans (and one placename) are claimed as from Berber. Two of them look acceptable, but all of them seem questionable, and they all refer to objects that there would have been no obvious reason to borrow terms for. It's possible that Berber influence can be found in southern Italian dialects, but this doesn't present a terribly convincing argument. Still, here they are:
  • źembr / źimbr / zimr / źimmr "billy-goat" (caprone, becco) < pan-Berber izimmər "ram", p. 39. (Looks good, but why the shift in species? - Also, see comments for an alternative Greek etymology.)
  • aččáta "big meal" (scorpacciata, mangiata, spanciata) < pan-Berber əčč "eat", p. 11. (The semantic and phonetic match are great, but the word is so short that coincidence seems hard to rule out.)
  • šéḍḍa "wing" (ala) < Zenati Berber "bird", eg Siwi ašṭiṭ, p. 26. The author mentions an alternative possibility - deriving it from Italian ascella "armpit" - that seems much more plausible.
  • Zaza (placename) < Berber azəzzu "thorny broom (plant sp.)" - not discussed in any detail (author cites Renisio), p. 41.

Saturday, March 07, 2009

Tawalt closing down

Tawalt is a nine-year-old Libya-focused Amazigh/Berber website with a remarkable collection of audio recordings, sketch grammars, vocabularies, and resources for some of the least well documented Berber languages - those of Tunisia, Libya, and Egypt. It is thus rather a shame that Tawalt is shutting down - updates stopping immediately, and site to go down by the end of the year. Sure, the Wayback Machine should preserve all the texts on it - but not its remarkable audio archives (which have already disappeared from the main page.) Their plans are probably related to political problems - the site's political postings had gotten rather outspoken. If you have any interest in Berber linguistics, I suggest looking around now before it disappears...

Wednesday, March 04, 2009

No, Berber isn't descended from Arabic

A few days ago I got lent a copy of a recent book in Arabic by Othmane Saadi: Dictionary of the Arabic Roots of Amazigh (Berber) Words معجم الجذور العربية للكلمات الأمازيغية (البربرية) (Tripoli: Academy of Arabic Language 2007.) My reaction, in brief, is that it's unscientific jingoistic claptrap. But I happen to have friends (not linguists, of course) who take it seriously; and I am told that the author, a proud member of the Chaoui Berber Nememcha (Nmamša) tribe, genuinely believes his own theory. I will therefore try to explain as simply as possible where the book goes wrong.

His starting point is noting the existence of strong similarities between Arabic and Berber in the vocabulary and grammar (p. C: “90% of Amazigh Berber words are pure or Arabised Arabic, and the grammar of Berber agrees with the grammar of Arabic.”) This is substantially correct, and has been known for a long time (see, for example, Igor Diakonoff's Afrasian Languages, Moscow: Nauka 1988, or at a more basic level one of my first posts), except that 90% is a substantial exaggeration – many of the comparisons he puts forward are at best questionable, as will be seen below. But he claims that the explanation for these similarities is that Berber descends from Arabic. Not just Berber either, as he says on p. B: “The term Arabitic عروبية means the ancient Arabic languages which are wrongly called the Semitic languages and which branched out from the source language Arabic thousands of years ago, such as Babylonian, and Assyrian, and Akkadian, and Phoenician Canaanite, and Aramaic, and Himyaritic, and Sabaean, and Thamudic, and Lihyanite, and Ma'inic, and ancient Egyptian, and Berber, and others.” Linguists subscribe to a rather different explanation for the observed similarities: that Berber and Arabic (and all the other languages he listed, and many he doesn't list such as Hausa and Somali) are all descended from a single language, called for convenience Proto-Afroasiatic (Greenberg 1950), which was different (and probably about equally different) from any of them.

How would you choose between these two hypotheses? Well, if the original language was different from Arabic, then you would expect some original forms to have been lost in Arabic but kept in other languages. Oddly enough, Saadi himself gives evidence for exactly that: he links the Berber ur “not” to Akkadian ul (p. 12), and the Berber -as “to him/her” to Akkadian -šu (p. 12), and the Berber nəkk “I” to Ancient Egyptian ink and Akkadian 'anāku, none of which are attested in Arabic. Unless you believe that Akkadian and Berber each independently invented the same new forms, or that they are more closely related to each other than to Arabic – which Saadi (correctly) does not claim – you have to conclude that the common ancestor of Arabic and Berber included words like ur/ul for “not”, and 'anāku for “I”, and so on, and hence was different from what we know as Arabic, just as it was different from Berber.

So maybe this common ancestor was Arabic in a different sense: Saadi argues that it was originally spoken in Arabia, so Arabic would be the one language that stayed at home, and presumably got less affected by foreign influence. Unfortunately, he doesn't have much of a case. His first argument (p. 1) is frankly risible: “Europe and North Africa were covered with ice before [18000 BC], whereas the Arabian peninsula enjoyed a climate similar to that of southern Europe now. The ice melted in the former and drought hit the latter, so mankind left the Arabian peninsula and settled North Africa and southern Europe.” The quote he cites on this actually says nothing about North Africa, and for good reason: even at the last glacial maximum North Africa was never covered by ice (see map), and was if anything more habitable before 18000 BC than it is now. He also notes (p. 2) that Berber princes have long claimed Yemenite origins. Such claims are questionable for many reasons (the desire for prestige, the originally matrilineal traditions of many Berber tribes, and no pre-Islamic attestations) – but even if true, it would prove nothing about the language: people change their language all the time without changing their ancestry, as any emigrant can tell you. The rest of his argument is a hotchpotch of miscellaneous quotes which at best claim that various early North African peoples or languages or cultures originated in the Middle East; in a particularly ludicrous case, he blithely quotes Bousquet (1957) to the effect that the Berber language “came from Asia Minor” [Turkey!] None of these quotes so much as mention the Arabian peninsula.

In fact, the linguistic evidence means that Proto-Semitic may well have been spoken in Arabia and certainly was spoken in the Middle East, but the common ancestor of Berber, Egyptian, and Semitic was most likely located in Africa. You see, as noted above, these three language families are also quite closely related to Chadic (spoken mainly in Nigeria and Chad) and Cushitic (spoken around the Horn of Africa) – which means that 4 out of 5 branches of this family are native to Africa. It is more likely that one branch left Africa than that 4 branches each separately followed the same narrow path across Sinai or crossed the Red Sea. (For theoretical background, see Campbell 2004.)

In other words: whether the similarities this book gathers between Arabic and Berber are valid or not, they don't do anything to support the author's claim that Berber descends from Arabic. Do they at least have the merit of being valid comparisons? Sometimes, but not with any consistency. Many of his comparisons look rather far-fetched, eg on p. D:

taməṭṭuṯ “woman” < Ar. ṭāmiṯ طامث “menstruator”
argaz “man” < Ar. rakīza(tu l-'usrā) ركيزة الأسرى “pillar (of the family)”
ixəf “head” < Ar. xf' خفأ “appear”, because the head stands out
tadaγt “armpit” < Ar. daγdaγah دغدغة “tickling”
alγəm “camel” < Ar. luγām لغام “the foam that comes out of camels' mouths”

Many others are clearly genuine loanwords, often featuring sounds that cannot be reconstructed for Proto-Berber, though I don't think many of these are original suggestions, eg:

(p. D) axərraz “cobbler” < Ar. xaraza خرز “to sew leather”
(p. H) abrid “road” < Ar. barīd بريد (confirmed by the Tuareg pronunciation of this word, abărid)
(p. 38) ləbṣəl “onion” < Ar. baṣal بصل (Siwi happens to preserve an older word for "onion": afəllu)
(p. 78) taħzamt “belt” < Ar. ħizām حزام

A couple are known Phoenician loanwords:

(p. 57) agadir, ažadir "wall" - Ar. jidār جدار

A few are well-known Afroasiatic cognates, and scattered among them may be other valid cognates:

(p. 250) iləs “tongue” - Ar. lisān لسان
(p. 110) iđammən “blood” - Ar. dam دم
(p. 292) tiqqad “burning” - Ar. wqd وقد

But the book makes no attempt to distinguish between words taken from Arabic comparatively recently and words inherited from the common ancestor of Berber and Arabic, and seems to assume that any word found in both dialectal Arabic (Darja) and Berber must automatically be originally Arabic, rather than possibly being a borrowing from Berber into Arabic. There is a well-known technique for sorting out inherited cognates from loanwords from coincidental similarities: sound correspondences. Sounds don't usually change at random: they change systematically, just as all j's in Egyptian Arabic become g. You establish which Berber sounds normally correspond to which Arabic ones under what circumstances, based on looking at what happens in the clearest cases; that gives you a standard by which to judge the doubtful ones. Saadi has made no effort to do this, and the unfortunate result is that in his comparisons the chaff far outweighs the wheat.

Berber and Arabic both descend from the same language, but that language was neither Berber nor Arabic, and probably didn't come from Arabia - and if you want to know about that common source, then you'll learn more from the works of Diakonoff or Greenberg, or even from more problematic sources like Orel and Stolbova 1999 or Militarev's online database, than from Saadi 2007.

Wednesday, February 25, 2009

Endangered Languages Week

It's half-over already, but I really ought to mention: Endangered Languages Week is happening at SOAS this week, and may be of interest to readers in London.

Also, a interesting news story, a reminder that many countries still have legal restrictions on what language you can speak where: A prominent Kurdish lawmaker gave a speech in his native Kurdish in Turkey’s Parliament on Tuesday, breaking taboos and also the law in Turkey.

Sunday, February 22, 2009

`baskundza igwạḍən!

I don't suppose there are more than about two or three people on earth who care, but I just figured out an etymology that's been puzzling me for a while. In Kwaṛandzyəy, the word for "genie" is agwəḍ, plural igwạḍən. It looks Berber for its form alone, but I had never found it in any dictionary - until now, going through Taine-Cheikh's new Zenaga dictionary, when I came across ugṛuđ̣an (original singular *ugṛuḍ) "démons, diables (plus dangereux, plus forts que les autres)". It turns out to have been borrowed into Hassaniya too - īgṛäwṭən. The loss of ṛ is more or less regular in Kwarandzyey (usually it's restricted to intervocalic positions, but there are a few other examples like this); so is the shortening of a long vowel to ə in a final closed syllable, with a w remaining to indicate its former quality. Quite possibly the next commenter will tell me that actually this word is well-known in Kabylie or Morocco or something, but for now it's another piece of evidence for my claim that Kwarandzyey includes a number of loanwords specifically from the Zenaga branch of Berber.

UPDATE: see comments - it wasn't the next commented, but the third one who established that this word is attested in southern Morocco too, which makes sense both since that region is also fairly close to Tabelbala and since it tends to be easier to find Zenaga cognates there than further north or east.

Friday, February 20, 2009

The Tyranny of Morphology

Coming out of an airport, you have to pick one of two exits: "Goods to Declare" or "Nothing to Declare". You have to go through one to get out; but (at least in Customs' eyes), by going through either exit, you state whether or not the contents of your luggage are legally subject to import duties. If you feel so scrupulously honest and so intensely secretive that you decide you have to leave that question unanswered - your only option is to stay inside.

Often your language does that too (Whorf said it first.) Just like the airports, the trick is to set things up in such a way that trying not to answer the question is either unacceptable (ungrammatical) or automatically interpreted as implying a particular answer. If you're talking about a friend in English, you don't have to indicate whether the friend is male or female until you refer back to the friend with "he" or "she"; in Arabic or Spanish, you have to state which it is from the start; and in Chinese or Songhay you can get away with never saying it at all. If you believe something definitely happens at some point, but don't want to say whether it's already happened or not yet, there's no simple way to say that. At best, you end up having to use cumbersome disjunctions like, if you're into apocalyptic prophecies, "The Antichrist either will be born some day or already has been"; and disjunctions like that will always be interpreted as meaning that you don't know which, not that you know but don't feel it's relevant.

In Korean (according to a talk by Peter Sells I heard today), a special verbal affix -si- (one among many, many politeness indicators) is used to indicate that the human subject of the verb (loosely speaking - it may also be a possessor of the subject, or a topic) is notionally of higher social status than the speaker. Thus:

sensayng-nim-i ka-si-ess-ta
teacher-HONORIFIC-NOMINATIVE go-SUBJECT.HONORIFIC-PAST-DECLARATIVE
"The teacher went."

vs.

koyangi-i ka-ess-ta
cat-NOMINATIVE go-PAST-DECLARATIVE
"The cat went."

The thing is, this means you can't be neutral about the subject. If you don't use this suffix with a subject that would normally take it, like "teacher" or "pastor", your listener will assume that you don't respect them so highly. You can't even get away with being ambiguous - I'm told that a disjunction of politeness levels, like *"The teacher went(honorific) or went(unmarked) away", is totally unacceptable. There are genres, such as academic writing or journalism, where politeness morphology is not normally used, allowing you to be neutral on this; but in a face-to-face conversation, as far as I understand, no such solution is available. (Any Korean readers should feel free to correct me!)

No language is likely to be able to stop you from saying what you want to say, if you try hard enough. But things like this can make it a lot harder to avoid saying what you don't necessarily want to say.

"Written in Islamic"

I don't usually do current events posts, but this one was cute enough to warrant a micro-post: egregious ex-Senator Rick Santorum declares that Muslims think that “The Quran is perfect just the way it is, that’s why it is only written in Islamic.” In most speeches, a sentence like that would be a major embarrassment; in this one, it's merely his only linguistics-related blooper.

(Via Angry Arab.)

Sunday, February 15, 2009

Fusha: the Straussian choice?

I came across a review of a book called Why are the Arabs not Free? The Politics of Writing, by an Egyptian psychoanalyst. I haven't read it (nor Adonis, whom he discusses below) but the quote presents an interesting perspective on Arabic diglossia:
My understanding of the political significance of this divorce between political and demotic Arabic and the key place of writing in the perpetuation of despotism crystallised when I read the work of our great poet Adonis, entitled The Book. It is one of the most revolutionary books I've read in Arabic literature. Apart from its provocative title, it lays bare the truth of our political history as having been a series of assassinations in a struggle for power. But it's written in such a high style that it's a difficult text even for the educated, without taking into account the vast majority of illiterate folk. So, it's no wonder that The Book has remained a 'dead letter'. I may say that I once heard Adonis declare that he won't ever write except in 'grammatical' Arabic because he prefers writing in a 'dead language'. One may wonder if his choice doesn't also represent his method for dealing with the condition [the German-born American political philosopher] Leo Strauss describes in his Persecution and the Art of Writing. The authorities are happy to ignore such books because in the unlikely event that they themselves have understood them, they know that their message will only reach a very limited number of people.
A tempting hypothesis in some ways, this idea that Fusha acts to insulate the majority of the population from the debates of intellectuals, keeping the powers that be safer from ideologically-inspired opposition and the intellectuals themselves safer (in the short term!) from popular reactions to their speculations. But is the issue really that people have trouble with the language, or just don't read much? Both are true to some degree, but in an era where TV shows and news programs in standard Arabic command large audiences across the Arab world, it's not plausible to blame everything on the difficulty of the language.

Elsewhere in the article he is said to imply that giving the colloquial greater status will "reduce any feeling of powerlessness as a result of a lack of formal linguistic expertise". That seems harder to argue with, given that many (probably most) people who can understand standard Arabic fine can't put together more than a sentence or two without mistakes, and certainly can't sound as eloquent or clear or at ease in it as in their colloquial language. But then again, what power does speaking standard Arabic well actually entail, when plenty of ministers and millionaires can't? Only the power to take part in debates that seem to have remarkably little effect on the society around them?

Friday, February 13, 2009

Why do historical linguistics?

Unraveling the details of a given language family's history is painstaking, detail-oriented work - comparing hundreds or thousands of words to each other, looking through different languages' grammars, coming up with hypotheses to explain what you see and hoping the next language you look at doesn't disprove them... Why do it?

Well, for one thing, you end up showing interesting things about the history of the relevant part of the world, often things it would be hard or impossible to show any other way - that Madagascar was settled by people from Borneo, for example, or that Ijo slaves from Nigeria ended up on the Berbice River in Guyana, or that Persians and Swedes (along with a lot of other people!) ultimately both got their language from a common source. But that depends on your being interested in a particular region; why would a person working on the historical linguistics of (say) the Sahara care about the historical linguistics of New Guinea, or Alaska, or even Europe?

It's because people are pretty similar everywhere - we all have roughly the same mouths and the same brains, and as a result we all tend to make roughly the same kinds of changes. Looking at changes in the languages of Europe, and at which direction they went, turns out to give you a pretty good idea of what kind of changes to expect in New Guinea - and vice versa; wherever you go, k is much more likely to change to g than to n, and a word meaning "want" is much more likely to become a future tense marker than a word meaning "jump".

That means that all these individual small-scale studies are so many pieces fitting together to form a map of how language works. Describing a language (no mean challenge in itself) shows you one set of possibilities; typology tells you the possible states of a language; but historical linguistics relates them to one another, showing you which states are closely linked and which are not. You can't predict what will happen to a language, but you can see in advance what kind of changes are likely and what kind are unlikely.

For sounds, this map of changes - this network linking different states of a language to one another - will seem familiar; it corresponds closely to articulatory and/or auditory similarity. You can mostly account for it by knowing how different sounds are made (with the lips, the tongue, etc...) and which sounds are hardest to distinguish. The key test for a theory of syntax (as far as I'm concerned) is whether it can account similarly for the attested map of syntactic change.

Thursday, January 22, 2009

Oldest Papuan writing?

What are the oldest written documents in a Papuan language (ie a non-Austronesian language of the New Guinea region?) I'm not totally sure, but a strong candidate has to be the court records of Ternate. The islands of Ternate and Tidore in eastern Indonesia speak two closely related languages belonging to the non-Austronesian North Halmaheran family. They have been writing using the Jawi Arabic script since at least the 1500s; in fact, some of the earliest surviving Malay manuscripts are letters from the sultan of Ternate from about 1521.

Recently I came across an 1890 book on Ternate online: Ternate: The Residency and its Sultanate. The book includes a brief introduction to the language and a word list; it also gives reproductions of several manuscripts whose originals date back to the mid-1800s, along with translations. So if you want to try your hand at deciphering them, or just see what a Papuan language looks like in Arabic script, have a look! The page I've linked to (Arabic interpolation de-italicised) starts:

ma-dero toma hijratu-nnabiyy ṣallī `alayhi wa-sallim nyonyohi pariama calamoi si-raturomdidi si-nyagisio si-rara, tahun alif, toma-arah Sawal, i-fani futu nyagimoi si-tomodi, malam Jumaatu...

"In the year Alif of the Moslem era 1296, during the month of Sawal, on a Thursday night, the seventeenth night of the moon..."

Wednesday, January 21, 2009

Verbal adjectives in English

It may seem pretty exotic to English-speakers that in some languages adjectives behave almost exactly like verbs, but this strategy is not as un-English as it looks. Consider the following colloquial American English sentences, with more formal approximate "translations":

This rocks! - This is good. (and not *This is rocking!)
That would rock! - That would be good.
That rocked! - That was good.
That was a rockin' day. - That was a good day.

This sucks! - This is bad. (and not *This is sucking!
That would suck! - That would be bad.
That sucked! - That was bad.
That was a sucky day. - That was a bad day.

In these, a property usually expressed with an adjective ("good", "bad") is being expressed using a stative verb, but only in predicative constructions (that is, to form a sentence.) In attributive function (that is, modifying a noun) an adjective derived from this verb is used. Like other stative verbs ("know", "be") but unlike non-stative verbs, it uses the simple present form to express a current situation, not the present continuous.

Within English, this pattern may seem pretty odd. But it corresponds rather well to how adjectives are expressed in Songhay languages, eg Koyra Chiini (Heath 1999:73). There, properties are expressed in predicative contexts just like verbs, with the same mood/aspect/negation particles, and in attributive contexts usually take a suffix:

ni beer - you are big (like ni koy - you went)
hal a ma beer - until it gets big (like a ma koy - he will go)
har beer - a big man

ni futu - you are bad
har futu-nte - a bad man

In Songhay the perfect aspect is used with stative verbs to express a current situation; but, like the English simple present tense, this is the simplest indicative verb form. The chief difference is that in Songhay the predicative verbs are used for inchoative senses too, as if "That rocks!" could mean "That is becoming good" as well as "That is good".

Typologically, I find it kind of interesting that what looks like a couple of verbal adjectives should be lurking in the recesses of the English lexicon. But it also has practical applications: if I were trying to teach Songhay or a typologically similar language to Americans, I would certainly start by discussing the example of these two English words.

Friday, January 16, 2009

Coptic adjectives

A little follow-up on the previous post, based mainly on Reintges' Coptic Egyptian (Sahidic Dialect): A Learner's Grammar:

In Coptic, predication of properties is handled exactly as for nouns, including the use of an determiner with the adjective:

hen-noc gar ne neu-polytia
indef.pl-great for are their-labours.
For their labours are great.

In attribution, the structure is Determiner - A - n - B, where A can be the noun and B the adjective, or vice versa:

ou-kohi n-soouhs: a-small n convent
t-parthenos n-sabê: the-virgin n prudent

To express the material of which something is made, you use the same structure, except that only B can be the material:

t-kloole n-ouein: the-cloud n light "the cloud of light"

Note that this is separate from the attributive construction:

ntof pe-iôt pahôm "He, our father Pahom"

So can adjectives be distinguished as a separate word class, when they behave so much like nouns? The answer is yes: an adjective is an item that can occupy either A or B in the attributive structure without a change in referential meaning. (See Coptic Grammatical Categories, Shisha-Halevy, p. 53.) If you reverse the constituents of a genitive or material construction, you change the referential meaning: "a vessel of wood" vs. "vessel wood (ie wood for vessels.)" If you do so for an adjective-noun attributive construction, the referential meaning stays the same: ou-noc n-polis or ou-polis n-noc both refer to the same entity, "a big city". So for this case, Dixon's hypothesis scrapes through.

Saturday, January 10, 2009

Adjectives - who needs 'em?

Most languages have a class of words that express properties and behave differently from other words. These are called adjectives. In English, for example, words like "red" or "old" or "tall" behave differently from nouns or verbs. For example, you add -s to verbs in the present tense if their subject is 3rd person singular, like "he sings" or "she eats"; but you can't add -s to an adjective, so you say "he is red" rather than *"he reds". You can put "very" before an adjective ("very red"), but not usually before a noun (you can't say *"very food".) Verbs can't be placed between "the" and the noun (unless you add an ending like -ing or -ed), but adjectives can (you can say "the red car", but not "the move car").

It turns out, according to Dixon 2004, that practically every language - perhaps every language - has at least one separate class of words, definable purely on the grounds of their (morphosyntactic) behaviour rather than their meaning, that refer to properties. This class typically includes words expressing size, age, value, and colour, and sometimes more.

But often, a concept expressed using an adjective in one language is expressed only by a verb or a noun in another. For example, in Kwarandzyəy adjectives come between the noun and the plural marker:

ạdṛạ kədda yu
mountain small PL
"little mountains" (hills)

But there is no adjective "happy" in Kwarandzyəy; instead, you use a verb, yəfṛəħ "be happy, rejoice". And to say "the happy people", you say "the people who are happy/have rejoiced":

bạ γ i-ba-yəfṛəħ
person who they-PF-happy

Moreover, though they may always be distinguishable by some test, they usually tend to behave very much like another word class. In fact, Stassen 1997:30 (link goes to 2003) postulates that in every languages adjectives handle predication (saying "X is red", for example) in the same way as either verbs, nouns, or locations. For example, in English or Arabic, adjectives handle predication like nouns (you say "He is tall", just like "He is a footballer"); in Korean or Tamasheq, they do it like verbs; and some languages, like Japanese, have both verb-like and noun-like adjectives.

So clearly people can do without some adjectives, and clearly the behaviour of adjectives tends to be very similar to the behaviour of some other word class. Why not do without them altogether? It would be easy enough to construct a language where no morphological or syntactic tests could distinguish adjectives from verbs, or from nouns. So if practically every language does take the trouble to distinguish them, there must be some pretty powerful cognitive motivation for it - and some pretty powerful historical tendencies acting to separate adjectives from verbs and/or nouns. The question isn't directly relevant to my current work, but it's worth thinking about.

Sunday, December 28, 2008

Siwa and its significance for Arabic dialectology

Hope all my readers are having/have had a great holiday.

A paper of mine, "Siwa and its significance for Arabic dialectology", should (inshallah) be appearing in ZAL soon-ish. Basically, there's a whole lot of Arabic influence on Siwi, including things you wouldn't expect to be borrowed, like Arabic's rather unusual method of forming comparatives from adjectives. However, this influence shows clear signs of deriving, not from any dialect currently used in or even particularly near Siwa, but rather from a more archaic one, with some resemblance to the dialects of other Egyptian oases quite distant from it and some features not attested in any other Arabic dialect of Egypt or Libya. In the 1100s, according to al-Idrisi, Siwa was inhabited both by Berbers and by sedentary Arabs; I suspect that the Arabs got assimilated into the larger Berber community and that much of the Arabic element of Siwi derives from their now-extinct dialect. If this sort of thing interests you, have a look (you can download it from the link at the beginning of this paragraph) and please feel free to comment on it here or by email.

Sunday, October 12, 2008

Tifinagh at Leiden

There were two more talks at Leiden that I should have mentioned, on a subject I've always been interested in - Berber writing systems.

Ramada Elghamis is working on a thesis about Tuareg writing systems, and described the purpose of "ligatures" (a more appropriate term would be "conjuncts") in the Tifinagh of the Air region of Niger. Tuareg Tifinagh allows a number of letter pairs (rt, zt, nk...) to be combined into a single letter. It turns out that this is not artistic license, but an essential feature of the script. In traditional Tifinagh, no vowels are written - but if two letters are combined into a ligature, that means that there is no vowel between them, thus resolving a lot of ambiguities. For example (from memory, so details may be wrong), t-m-r-t is read "tamarit", a woman who is loved, whereas t-m-rt is read "tamart", beard; in unvocalised Arabic script, or in traditional Tifinagh minus the ligatures, there would be no way to distinguish the two.

Robert Kerr came up with a nice argument that Libyco-Berber, the pre-Roman script from which Tifinagh is descended, was adapted specifically from the Punic (early Carthaginian) variant of the Phoenician script, not the original Lebanese one and not the later Neo-Punic one. Basically, Old Phoenician marks no vowels at all; Punic marks a few vowels, almost always final ones; and Neo-Punic marks most vowels in all positions. Libyco-Berber (and traditional Tifinagh) also marks vowels only in final position; this rather odd idiosyncrasy is best interpreted as having been adopted from Punic rather than independently innovated.

Friday, October 10, 2008

Berberologie colloquium at Leiden

I've spent the past couple of days at the Berberologie colloquium in Leiden, and it's been great fun. There were plenty of very interesting speakers, but for me two languages stole the show: Tetserrét and Ghomara.

Tetserrét (discussed by Cécile Lux) is spoken by a Tuareg tribe, the Ayt-Tawari, in Niger. But it's not linguistically Tuareg at all - its closest relative is Zenaga, the Berber of Mauritania (not northern Berber, contrary to Wikipedia), and Tuaregs can't even understand it. It seems to be an isolated survival of the Berber language spoken in the region before the Tuareg got there. It's not in Ethnologue either. (Taine-Cheikh's new Zenaga dictionary is out, by the way, and was selling as fast as a book reasonably can in a conference of twenty people.)

But Ghomara, in northern Morocco, is something else. Across Berber, borrowed Arabic nouns typically behave like in Arabic (keeping their Arabic plurals, and not changing for case.) In Ghomara (discussed by Jamal El Hannouche), Arabic adjectives take Arabic rather than Berber agreement marking - and even some Arabic verbs get conjugated fully in Arabic, not in chance code-switching but regularly by all speakers, and up to and including pronominal object suffixes. It's not quite unprecedented worldwide, but that level of contact influence is pretty darn rare.

I didn't put Tadaksahak in the first paragraph because it's much less unfamiliar to me, but Regula Christiansen's paper on that had some interesting implications. Basically, Tadaksahak has all but lost the Songhay method of forming attributive adjectives; instead, it's substituted a simplified version of the Tuareg one (suffixing -an), which has become productive for Songhay adjectives too. The funny part is this: Songhay has a lot of CVC adjectives (stative verbs). Tuareg doesn't really do CVC adjectives; it prefers longer words. So when you add the -an to these, you typically reduplicate the adjective. For example, kan "be sweet" > kankanan "sweet". This comes worryingly close to invalidating a conjecture I had made on the borrowability of templatic morphology (but not quite!)

My own paper established that much of the Berber element of Kwarandzyey derives from an extinct close relative of Zenaga. In effect, the "Western Berber" genetic subgroup of Berber has four members: Zenaga itself (finally with a decent dictionary), Tetserrét (awaiting further publications), the large Berber element of Hassaniya, and part of the proportionally larger Berber element of Kwarandzyey.

Saturday, October 04, 2008

Translating from linguists' English to normal English

Machine translation between languages is hard, obviously. There are all sorts of reasons why just looking words up and constructing syntactic trees and changing orders appropriately isn't enough to produce a good output - mainly, the fact that to disambiguate ambiguities you often need real world knowledge, and different vocabularies are not always organised in the same way. How much that matters is really emphasised by thinking about a slightly different problem: translation from a technical vocabulary to a non-technical one within the same language.

Take the following sentences, pulled at random from a grammar on my shelf (Stroomer's Grammar of Boraana Oromo):
"Nouns ending in -ni (mostly -aani) have ultimate or penultimate stress in free variation."

"Verbs with the verb extension -ad'd'-, -at- have an AFF.IMPER.sg: -ád'd'i, -ád'd'u and a NEG.IMPER.sg: -atín(n)i, see 10.10." (p. 72)

If you are, say, a foreign worker about to be posted to northern Kenya, or a second-generation emigrant Oromo planning to go back and visit, you may well want to try and learn some Oromo from this book. But the odds are you will not know what either of these English sentences means, and that applies to quite a lot of the book.

How could you translate these sentences into terms a wider audience would understand? If you can assume a certain amount of basic knowledge (traditional parts of speech, consonants and vowels) then that makes things easier:
"Nouns ending in -ni (mostly -aani) get stressed on the last or second-to-last vowel, it doesn't matter which."

"Verbs with -ad'd'-, -at- added at the end have an imperative singular: -ád'd'i, -ád'd'u and a negative imperative singular: -atín(n)i, see 10.10."
Realistically, you can't assume that level of knowledge, certainly not in Britain at any rate (I still can't believe that what little grammar gets taught in schools here only ever seems to get taught in foreign language classes, not in English ones; that no doubt explains part of the country's comparatively low foreign language skills.) So what does that leave you with? Something like:
"When you say a word that refers to a person, place, or thing* and ends in -ni (mostly -aani), you put the emphasis at the end or just before the end, it doesn't matter which."

"If you have a word that means doing something* that has -ad'd'-, -at- added at the end, then to order one person to do that you add -ád'd'i, -ád'd'u, and to order them not to do that you add -atín(n)i, see 10.10."
(*Yes, I know that syntactic tests like whether they can be the object of a preposition yield more accurate definitions, but in practice these are a good first approximation, and the former does work even on gerunds: "Killing is a bad thing", so "killing" is a noun, but *"Kill is a bad thing", so "kill" isn't.)

Could this be done algorithmically? A simple substitution table would certainly not be enough. Just try it with any set of definitions you can think of:
"Words referring to a person, place, or thing ending in -ni (mostly -aani) have final or pre-final emphasis such that it doesn't matter which."

"Words that mean doing something with the words that mean doing something extension -ad'd'-, -at- have an agreeing order-giving one-entity: -ád'd'i, -ád'd'u and a denying order-giving one-entity: -atín(n)i, see 10.10." (p. 72)
Not terribly helpful, I think you'll agree... To come up with something a little more helpful (and I'm sure my renditions could be improved on) we had to change the whole structure of the sentence. Even then, at some point it's probably going to be more effective to just teach the person the grammatical notions and let them go forward from there than to keep giving brief explanations of the same notion over and over again.

The problem is certainly not unique to linguistics. Medicine, law, ecology - most fields have technical vocabularies that pose an obstacle to non-specialists, who will often have good reason to be interested in trying to make sense of them. Is there any role for algorithms in this (apart from obvious things like hyperlinking technical terms to dictionary entries)? It's well outside my usual field, but it would be interesting to hear of any attempts.

Saturday, September 13, 2008

Overheard from the code-switching department...

...from an Algerian here in London:


kanu supplying-lna
they.were supplying-to.us


You have a non-finite English form ("supplying") in a past continuous form, in accordance with the English construction but contrary to the Algerian Arabic one, which would require a finite form ("they supply"). You have an Algerian Arabic clitic pronoun - a form that can't stand on its own, but has to be attached to the end of something else - being stuck onto a totally unadapted verb in another language; code-switching in the middle of a phonological word! The facility with which some Algerian long-term residents of the UK combine their two languages is really rather remarkable, and would merit further study.

Thursday, September 11, 2008

Fieldwork and address books

Linguistics, with its regular sound shifts, unidirectional grammaticalisation processes, and tree diagrams, is perhaps the most satisfyingly scientific of the social sciences. But today I found myself reminded that it is still emphatically social, particularly when you want to actually gather new data about undocumented languages. Mobile phones have become ubiquitous even in such far-flung corners of the Sahara as Tabelbala and Siwa, used even by illiterate people - making it possible to keep asking people about the language well after you've gotten back to the university. So over the past months of fieldwork my phone has accumulated quite a lot of numbers, which I backed up to my computer today. The final count? At least 84 phone numbers from Tabelbala and 43 from Siwa. To put this in perspective, there are only about 3000 Kwarandzyey speakers, so I can call something like 3% of the population.

The field linguistics courses at SOAS lay a commendable emphasis on teaching the practicalities of fieldwork - what microphone, what recorder, what software... But there's a gap in the course: managing contacts. Going through these I found a few casual contacts I could barely or even not at all remember, and some people I could remember but not easily remember the relationships between. There's some information in my field notebooks, but it's scattered and not always detailed. I should have been making concise but informative notes about all these people somewhere as I took their numbers - not something you can do easily with my already somewhat antiquated mobile, but that might be a reason in itself to take a more sophisticated one along, or even to use a paper address book, if you have space in your pocket for one alongside your field notebook. If you plan to do any fieldwork, bear this in mind!

Thursday, September 04, 2008

Desert lizards


If you're an Arabic speaker from the right part of southwestern Algeria, you probably call the smooth-skinned sand-burrowing lizard referred to in English as "skink" šəṛšmala شرشمالة. I recently found the original form of this word in Al-Hilali's Berber-Arabic lexicon from 1665: asmrkal or asrmkal أسرمكال, a word composed from asrm "worm" and akal "earth". In many Berber varieties (the so-called Zenati ones), akal becomes šal, and in some Arabic dialects if there's one š ش in a word any s's س have to become ش, so you'd get شرمشال, and by metathesis شرشمال.

Are any readers familiar with skinks? What would you call them?

Wednesday, August 20, 2008

Triliterals in strange places

In a grammar I was looking at lately, I came across the following sentences:

"Nouns may be verbalized, or verbs nominalized, simply by bringing the stem into a suitable rhythmic form... Most of the rhythmic patterns call for a tri-consonantal stem. If a stem is di-consonantal in its primary form, a consonant (usually the glottal stop) is added to give it the proper structure... Often in the course of forming derivatives, stems that are too long are forced into one or the other of the regular patterns. They are cut down by the loss of quantity or of vowels or consonants as may be necessary."

Was this a Semitic language, or perhaps some less well known Afro-Asiatic cousin? No: this was Sierra Miwok, the pre-conquest language spoken by the Native Americans of central California inland from the Bay. (See map.) The "rhythmic patterns" only involve changes in quantity and CV>VC metathesis, not insertion of specific vowels as in Semitic, but the parallel is striking. Here are a few examples:

leppa- "to finish", with a CVCVCC pattern imposed, becomes lepa''- (gaining a glottal stop).
ṯolookošu- "three", with a CVCCV pattern imposed, becomes ṯolko- (losing the š).

Compare Arabic:
'ab- "father", with a 'aCCaaC plural template imposed, becomes 'aabaa'- "fathers", gaining a glottal stop (historically a semivowel, but never mind that)
`ankabuut- "spider", with a CaCaaCiC plural template imposed, becomes `anaakib- "spiders", losing the t.

Reference:
Freeland, L. S. 1951. Language of the Sierra Miwok. Baltimore: Waverley Press.

Tuesday, August 05, 2008

Nepal's language riots

Qatar is of course one of the most multicultural places on earth - citizens are only a small minority of the population, and even they include a lot of pre-oil era immigrants from Asia and Africa. Among the largest national groups here in recent years is Nepalis, so it's no wonder that the papers here in Doha have been full of a language controversy that readers elsewhere may not have noticed - the anti-Hindi riots in Nepal.

Apparently, the people of the plains in southern Nepal have ethnic ties to India. They don't speak Hindi natively, but commonly use it as a lingua franca between them. The new vice-president Parmanand Jha comes from this region, and decided to take his oath of office in Hindi (although his native tongue is Maithili). Highlanders took this as a deliberate snub to the official language Nepali, the worse for having not even been in his own language but rather in one primarily associated with India - and a week or so of riots, in which at least 10 people were injured, followed. He issued a sort of apology that calmed things down, but apparently now there are fresh protests from a plains group without Indian ties, the Tharu.

Those who prefer a jargon-filled angle on all this can regard this as an interesting case study in the symbolic weight of language choice in a multilingual context. In this case, they seem to have at least two different diglossias going on: Nepali vs. others in the hills, Hindi vs. others in northern India, and both languages effectively trying to claim the role of the high-prestige language in the plains in between, through competing political parties. Kind of reminds me of North Africa, actually... (And that's without even getting into the role of English.)

For a few links, try:
OhMyNews
KantipurOnline
Hindustan Times

Wednesday, July 30, 2008

Some surprising language links

Sorry about the infrequent posting, everyone - I've been burying myself in transcribing a few of my field recordings. There's plenty of interesting stuff on them: what to sing to encourage locusts to go away, how to make tea the proper Saharan way, tickling rhymes for kids... So naturally today I'll post a potentially linguistically interesting audio link: Library of Congress: American Memory Sound Recordings. This has a bunch of interviews with ex-slaves from the 1930s, which I understand from reading John McWhorter's Defining Creole have revolutionised Creolists' ideas about the history of African-American English. A popular theory had argued that during the era of slavery African-Americans spoke a creole English much like the ones in the Caribbean; these recordings, which show a speech not too different from today, turned out to pose a severe challenge to this view. It also has an Omaha pow-wow and late 1930's Californian folk music in a surprising number of languages, including Armenian, Finnish, and Gaelic. Look around.

Friday, July 11, 2008

How not to make an official Amazigh webpage

As I just learned from a comment on Awal nu Shawi, Algeria's High Commission for Amazighity, the government body set up in 1995 in response to the demands of Amazigh (Berber) activists for cultural recognition, finally has a webpage. I wish I could even manage to be surprised that the site is written exclusively in French - the only Tamazight present is in the title, and they don't have even a single word in Arabic - or that the content is meager and largely bureaucratic. How is it that our next door neighbour Morocco - with a rather less militant Amazigh movement and a rather smaller budget - can manage a beautiful, trilingual, fairly content-rich website like IRCAM for its equivalent body (and put the page up much earlier, at that), while Algeria's own HCA can't even be bothered to translate its website into a single Algerian language? Can you imagine going to the Academy of the Arabic Language site, say, and finding the whole page in French? A site like this makes it seem like its producers are interested neither in promoting the development of Tamazight nor in communicating with the majority of Algerians who read Arabic better than French. The Amazigh movement in Algeria is frequently accused of being just a Trojan horse for the promotion of French language and culture; you would think the HCA would take more trouble to avoid seeming to confirm this accusation.

Monday, July 07, 2008

One word, two masters: demonstrative agreement with addressee

In Qur'anic Arabic (this is hardly ever applied in Modern Standard), at least in presentative contexts, the word "that" agrees in number and gender not only with the noun it refers to, but also with the addressee. (A YouTube video lecture on this by some shaykh is available for Arabic-speakers.) "That" is morphologically composed of two elements. The first bit agrees with the referent:

đā-li-masculine singular
ti-l-feminine singular
đāni-masculine dual
tāni-feminine dual
'ulā'i-masculine/feminine animate plural


The second bit agrees with the addressee:
-kamasculine/feminine singular
-kumāmasculine/feminine dual
-kummasculine plural
-kunnafeminine animate plural


(In modern standard Arabic, only -ka is normally used here; even in Qur'anic contexts, the other forms' usage seems to be limited.)

Thus in Surat Yusuf, verse 32, Pharaoh's wife, addressing her women friends, says:
فَذَلِكُنَّ الَّذِي لُمْتُنَّنِي فِيهِ
fa-đālikunna llađī lumtunnanī fīhi
That is the man about whom you blamed me!

Then in verse 37, Yūsuf/Joseph, addressing his two cellmates, says:
ذَلِكُمَا مِمَّا عَلَّمَنِي رَبِّي
đālikumā mimmā `allamanī rabbī
That is (part) of what my Lord has taught me.

Likewise, in Surat al-A`rāf 22, God tells Adam and Eve:
أَلَمْ أَنْهَكُمَا عَن تِلْكُمَا الشَّجَرَةِ
'a-lam 'anhākumā `an tilkumā ššajarati?
Did I not forbid you from that tree?

It's not hard to come up with a story for how this grammatical phenomenon could have emerged: li in Arabic means "to" or "for", and the endings it takes are (with one exception) the same as those above, so it could easily have either conveyed a presentative meaning (compare English expressions like "That's London for you!") or, less probably in this case, indicated proximity to the addressee ("Get me that book next to you"). But I have a question for all you wonderful readers who have gotten this far: do you know of any other language that does something like this?

Friday, July 04, 2008

The Berbers of southern Egypt

Checking through the 11th-century geographer al-Bakri for information on the linguistic history of Siwa, I was not surprised to see that he says the Siwans were Berber, and not very surprised to see that the people of Bahariya at the time were Arabs and Copts, and those of Farafra Copts alone. I was a bit more surprised, though, when a little further down the page he says that some of the people of Dakhla and Kharja, in southern Egypt (map), were Lawāta Berbers:

وهذا واح الداخلة كثير الأنهار والعمارات... ومن هذا الواح إلى الواحين الخارجين ثلاث مراحل وهو آخر بلاد الإسلام... وفي بغض الواحات قبائل من لواتة


It kind of fits with an observation made by several Nubian specialists (I read it in an article by Bechhaus-Gerst) that Nubian - specifically Nobiin, in fact, not the Nubian languages of Kordofan or Darfur - seems to contain Berber loanwords; the easiest to remember, and most convincing, of these is "water", aman (which in other Nubian languages is something completely different, along the lines of essi.) If a dictionary of the Arabic dialects of these oases ever comes out, it would be very interesting to check it for Berber loanwords.

On a more romantic note, al-Bakri also warns those travelling into the desolate lands west of these oases that they will find "great sands... full of palm trees and springs, with no civilisation nor companions, where the murmuring of the jinn is heard unceasingly."

Friday, June 13, 2008

This Post is a Sin to Read

I imagine pretty much all English speakers agree on the grammaticality of the following sentence:
* It is a sin to eat pork.

But looking around online recently, I was struck by the following construction:
* Pork is a sin to eat
* Soon it will say in the bible that Speghetti is a sin to eat.
* I don` t think any kind of food is a sin to eat

To me, this construction seems rather odd, and the extreme rarity of such constructions on Google suggests that I'm with the majority of English speakers on this point. Do people who do find this normal allow it with other verbs, I wonder? Can they say "This post is a sin to read?" or "Wine is a sin to drink?" Or, indeed, "Tea is a pleasure to drink?" Has anyone else heard constructions along these lines? Presumably, these speakers were influenced by the analogy of sentences like "A mind is a terrible thing to waste" or "Tea is a good thing to drink"; but if I ever figure out why the former seem so weird and the latter are perfectly grammatical, I'll make sure to tell you...

Tuesday, June 10, 2008

Baby talk across the centuries

Most languages probably have a few words used especially for addressing babies. However, Siwi seems to have a lot more than I know from English or Arabic (I've recorded something like 40). One of these (already noted in Laoust 1931) is mbuwwa "water" (the normal Siwi word is aman). mbuwwa, meaning "water" or "drink", turns out to be rather widespread: they use it in baby talk in Syria, Lebanon, Tunisia, Algeria, Morocco, Malta, Sicily, and probably a few other places for which I haven't found sources. The remarkable part is that Ferguson managed to track down a historical source for this word. Varro, a Roman grammarian of the first century BC, gives bua as the nursery word for "drink" (presumably to be related to bibere, the adult verb for "drink".) (Unfortunately, I haven't managed to find the relevant work online.) If the connection is correct, then this word (possibly along with some others, like pappa for "bread" or "food") has persisted in Mediterranean baby talk for at least 2000 years, apparently without ever passing into adult speech.

So what special words do you use in your language when talking to babies?

Sunday, June 01, 2008

Kant's Sparrow and the Wolf Girl

I remember coming back to Algeria after a year or so in America at the age of six. I had completely forgotten the Arabic I had known, and relearning it was an incredibly difficult process that took years, made no easier by my frequent preference for books over playmates. Over the past seven months, I've found learning Siwi and Korandje far, far easier than learning Arabic was then, and I'm pretty sure I speak both of them, if not fluently, at least far better than I spoke Arabic after my first four months back. Yet the nativist theory of language acquisition that I remember from my linguistics courses says that kids should learn languages much more easily than adults. I don't expect anybody to pick theories based on anecdotal evidence from my childhood memories, but this has made me wonder again whether kids usually learning languages faster and better is due to a pre-programmed critical period for language learning, or simply to the big difference between the social contexts of adults and children. Coincidentally (being back in London), I found two works vaguely relevant to that question this weekend; neither offers an answer, but they are interesting background.

Kant's Sparrow confirms a claim originally reported by Kant - that sparrows brought up by canaries learn to sing like canaries. Apparently, they do - but not completely. Not only do their canary songs feature a detectable accent (they differ in several ways, notably in repeating the same syllable fewer times in a row), but their repertoire includes some song types ("two-voice syllables") which they rarely or never heard from the canaries raising them, and which the author attributes to sparrows' innate repertoire (3.3.2.6.) In other words, sparrow song, like human communication, combines innate and learned (arbitrary, if you like) elements.

Wolf Child and Human Child, by Arnold Gesell, is a short, not very helpful work on a very interesting case, apparently described more fully in Diary of the Wolf Children of Midnapore, by Rev. J. A. L. Singh - two children, later named Kamala and Amala, who were adopted into a wolf family, and raised for years alongside the mother wolf's own cubs. In 1920, in response to locals' reports of a "man-ghost", a party of men dug into the wolf's den, killed the mother wolf when it tried to fight back, and brought the two children back to be taken to an orphanage (and the two wolf cubs they lived with to be sold at a fair.) Kamala was about eight, and Amala substantially younger; however, Amala died only a year later Unsurprisingly, Kamala found language rather difficult to acquire; even without the wolf factor, I imagine losing your entire family and then your entire step-family before the age of nine might have a negative effect. At any rate, apparently, she spoke her first word two years after being captured, and her first two-word sentence after three and a half years. For later years the information gets a lot sparser, but it is claimed that by the time the poor kid died (from illness) nine years later, at the estimated age of seventeen, she "talked freely with full sense of words used." The report that after several years "her formerly rigid countenance took on more expression" suggests a similar gradual development in her body language. However, while at eight years old she knew little or nothing of how humans communicate, she seems to have learned at least some wolf methods - for months at the orphanage, she would howl every night, at 10 pm, 1 am, and 3 am, and when approached by someone she did not trust she would show her teeth. Unfortunately, the lack of detail makes it hard to say what this says about first language acquisition - how well could she really speak before she died? Perhaps Rev. Singh's diary offers some quotes.

NB: see comments; apparently there is serious doubt about the veracity of this account. Looks like I should have Googled first.. The original diary also turns out to be online.

Thursday, May 22, 2008

African influence on native Nicaraguan languages!

...and I bet that got your attention, if you're the sort of person who reads this blog.

Ulwa is a language native to the eastern highlands of central Nicaragua, and now spoken mainly in Karawala on the Atlantic coast. It belongs to the small Misumalpan language family, along with Miskito; an interesting characteristic of this family is the position of nominal possessive affixes, which may be suffixed or infixed depending on the word's syllable structure. The Miskito kingdom had a longstanding relationship with the British, as a result of which English Creole is widely spoken on Nicaragua's Atlantic coast; both Miskito and English have influenced Ulwa, as has Spanish of course. You can find a nice dictionary and a brief grammar at the Ulwa Language Home Page.

Anyway, the Ulwa word for "east" turns out to be mâsara. I'm sure some readers will already be thinking of Maghrebi Arabic/Berber mâṣəṛ (from Arabic مِصْر), with reflexes in a variety of West African languages along the lines of masara - meaning Egypt! Unfortunately, a second glance reveals that "west" is mâ âwai, suggesting that maybe mâsara is some kind of compound with mâ. mâ, sure enough, turns out to mean "sun", while sara means "origin". So much for that idea; but what a good example of how a coincidental lookalike can emerge. I can't find any similar way to explain the word for "God", though - which is Alah...

So what about that African loanword I promised? There really is at least one, but it is somewhat less exciting. "Peanut", in Ulwa, is pinda. This word, referring to a post by Polyglot Vegetarian, appears to derive from Kikongo m-pinda, and was borrowed into English as pindar (various spellings) before being ousted by peanut. So this word may have been mediated by English, but is of clear Kikongo origin - sensibly enough, given that peanuts themselves come from Africa. If you want more African loanwords into Caribbean Native American languages, try Garifuna - where the word for "man" is a Bantu loanword.

Sunday, May 18, 2008

Ode to repression II

In response to mild popular demand, here's the original of the poem I translated in the last post, in Kabyle orthography for convenience, although this orthography doesn't fit Siwi perfectly - just remember that "ay" (or "a y", or "a i") is to be pronounced like French é. (For those not familiar with this system: "e" is a short schwa, "c" is sh, "ɛ" is Arabic `ayn.) Two points that may help for speakers of other Berber languages: in Siwi the negative is la (not ur), and the future is marked with ga (not ad).

kell ma qedṛaṭ kmec elbed,
la tac-as esserr i ḥedd
γayr belɛ-a netta la ikemmed
kan jebdaṭ-t af cal ga yebṛem
amra wenn ga iṣaṛ-ak ektem,
ejj-a γayr ṛebbwi ga yaɛlem

كلّ ما قدراط اكمش البد
لا تاشاس السّرّ إي حدّ
غير بلعا نتّا لا يكمّد
كان جبدات آف شال گا يبرم
آمرا ونّ گا يصاراك اكتم
اجّا غير ربي گا يعلم

In a village society where everyone knows everyone else and will still be neighbours with everyone else thirty or fifty years on, particularly one that puts a high value on keeping up appearances and presenting a good face to the world, there will always be a lot of thoughts and memories that are best kept to oneself for the sake of keeping one's relations with others good and one's public image unblemished - personal disagreements or dislikes, unfulfillable desires, actions that run counter to the social code... what Ernest Gellner used to call the tyranny of cousins rather than the tyranny of kings. That's what this poem is about: you may be in love with someone unavailable, or you may have reason to hate someone you're supposed to respect, or whatever, but you can't talk about it because of the scandal it would create and the negative impact that would have on yourself and your family. I suspect that if you've ever lived in such a place, you'll get the poem, and if you're born and bred in the city, you probably won't even with this explanation; but tell me if I'm wrong.

Wednesday, April 30, 2008

Ode to repression

No, not in the political sense, in the psychological one... Just thought I'd share a piece of an excellent Siwi poem that struck me as characteristically North African, with a theme reminding me strongly of Dahmane el Harrachi's song "Khabbi serrek yalghafel" (Hide your secret, neglectful one). Obviously, it doesn't work as well in my attempt at translation, but here goes:
Whatever you can, tie up and hide,
Don't give anyone a secret, on any side,
Just swallow it, it won't hurt inside.
If you let it out, it'll do the rounds.
Keep what happens to you underground,
By God alone to be finally found.

Wednesday, April 16, 2008

Update from Siwa

Hi everybody! I'm in Siwa, and things are going well. The oasis is so much bigger and more prosperous than Tabelbala it seems almost decadent by comparison; its lakes and its expanses of groves suggest some idea of what Tabelbala's environment might have been like at its peak. The language is in no immediate danger; while some words are disappearing due to the great change in lifestyle, not only do children all seem to speak Siwi as a first language, but a substantial portion of the Shihaybat Bedouin settled in the western edge of Siwa learn it as a second one. However, the declining popularity of music at weddings may to some degree be threatening the vigorous local tradition of Siwi-language poetry. As Vycichl noted, Siwi has grammatically conditioned stress; in fact, you could argue that case is marked in Siwi by stress shifts. Siwi is definitely not mutually comprehensible with Kabyle, by the way - I've now tested this in both directions - nor with any Moroccan variety, according to local watchers of Moroccan satellite channels. Gara is also an interesting place - a much poorer, smaller oasis a hundred-odd km off, inhabited by mainly black people speaking Siwi. I've been there, but unfortunately security regulations more or less preclude spending the night.

The Bedouin Arabic of western Egypt is also of some interest. It is remarkably conservative, though not as much so as the dialects of Najd - it has a fully productive dual, distinguishes masculine and feminine plurals (both for verbal and adjectival agreement), and still has most short vowels. Technically, it shares some of the defining innovations of Maghrebi Arabic, in particular the 1st person plural n-...-uu; but it sounds scarcely closer to Algerian than even Cairene Arabic. They write a lot of poetry, some of it rather good. Inconveniently but interestingly, it appears that most Arabic influence on Siwi derives neither from their dialect nor from Cairene.

On a final note, anyone interested in medieval Berber history (there must be someone...) will recall the rather large Huwwara tribe (from which Houari Boumedienne ultimately got his nom de guerre). It turns out they're still very much around in the western Delta and even Upper Egypt, although they all speak Arabic now, as they had already begun to do in Ibn Khaldun's time; I met a Huwwari just the other day.