Wednesday, April 27, 2011

An atom's weight of philology

One of the oldest motivations for studying the history of language is to better study the fixed texts of holy books or classics. We try to learn from such texts, but without an understanding of philology we misread them - because, while the words have remained the same, their content has changed. Ibn Quraysh is one case in point; Ruskin offers another:
"[I]n languages so mongrel of breed as the English, there is a fatal power of equivocation put into men's hands, almost whether they will or no, in being able to use Greek or Latin words for an idea when they want it to be awful [ie impressive]; and Saxon or otherwise common words when they want it to be vulgar… [C]onsider what effect has been produced on the English vulgar mind by the use of the sonorous Latin form "damn", in translating the Greek katakrínō, when people charitably wish to make it forcible; and the substitution of the temperate "condemn" for it, when they choose to keep it gentle; and what notable sermons have been preached by illiterate clergymen on - "He that believeth not shall be damned"; though they would shrink in horror from translating Heb. xi. 7, "The saving of his house, by which he damned the world"… "
Standard Arabic has no layer of prestige loanwords corresponding to Greek and Latin words in English - all the classics of the Arab world are themselves in Arabic, and great efforts have been expended to keep the grammar of Standard Arabic roughly constant since the pre-Islamic era. But, thanks to the many new meanings conferred upon old terms during episodes of massive translation - both in the modern era and the Abbasid era - it is fairly susceptible to another of Ruskin's complaints: misinterpreting the words of old texts thanks to their modern meanings.

Once a medical student at Cambridge told me in all seriousness that the Qur'ān anticipated modern science by centuries in mentioning the "atom" (فمن يعمل مثقال ذرة خيرا يره, "for he who does an atom's weight of good shall see it")! Of course, every modern educated Arab knows that a ذرة dharrah is an atom. But looking at a pre-modern dictionary, such as Lisān al-`Arab, gives a rather different picture: a dharrah then was a type of small red ant, a weight equivalent to 1/100 of a barley grain, or a mote of dust (as seen in sunbeams), not an elementary particle of which all matter is composed. In parts of Sudan the first of those meanings is still in regular use: dirr there means a type of ant. But elsewhere they all seem to have faded from away from popular speech.

If I were interested in an English word, I could easily look it up in the OED and find a complete history of its different meanings and the dates at which they were attested. But for Arabic no such dictionary exists; to figure out when and how dharrah came to mean "atom" in the modern sense, I would have to look through a bunch of pre-modern works, or find an article on the subject. It's a gap that would be well worth filling.

Thursday, April 14, 2011

Why *h1 and *h2 were not valid onsets in late proto-Berber

I've been working on my hopefully-forthcoming book about Siwi and thinking more about Berber laryngeals (see also Phoenix's recent post), two tasks that intermesh rather handily. Now Siwi has a wide range of strategies for forming the intensive (ie, in Siwi, the realis imperfective) of verbs, not obviously related to one another. But it is usually possible to predict which will be used from the form of the root. Basically, to recap the relevant page of my thesis, ignoring the fəl verbs discussed in the previous post and some other synchronic irregularities (U=consonant or full vowel; either count as a unit of the root):

Prefix t-:
- to geminate-initial roots
- to roots with the mediopassive prefix ən-
- to vowel-initial roots
- to vowel-medial (CVC) roots
Geminate U2:
- when U1 and U2 are distinct consonants, and U2/U3 is final
Put -a- after consonantal U3, changing any previous full vowels to a:
- when the last two units are distinct consonants (unless geminate-U2 / prefix-t applies), or
- when U2 is a full vowel (in which case prefix-t also applies)
Suffix -u:
- to geminate-final roots

Can we further simplify these conditions? In particular, what do the rather disparate environments to which t- is prefixed have in common?

Well, Siwi, like most Berber languages, shows the so-called “mobile schwa” phenomenon – ie, the position of schwa is mostly predictable solely from the consonants and long vowels of the word. (Basically, you put a schwa between any two adjacent consonants followed by a consonant or word boundary, starting from the left cyclically.) This also means that the coda/onset status of a given consonant in a stem is predictable, and depends on the affixes – for example, the k is a coda in əktər “bring!”, but an onset in kətr-ax “I brought”. However, there are a few exceptions to this principle – clusters that cannot be broken up by schwa, or, equivalently, codas that do not become onsets. These include:
- geminates: geminates cannot be broken up by schwa, and the first element of a geminate is always a coda.
- mediopassive ən-: the cluster ən+C that it forms cannot be broken up by schwa, and the n is always a coda (except before the borrowed voiced pharyngeal ʕ.)

Full vowels are by definition not onsets (semivowels behave quite differently from full vowels in Siwi.) So we can reduce the first three conditions for t- to a single one: t- is used when the first element of the root is not an acceptable onset. The fourth condition seems to be separate.

The use of t-/tt- under two of three of the conditions we have unified is reconstructible for proto-Berber (mediopassive ən-, or at least its syllabic structure, is a borrowing from Arabic), so it would be reasonable to reconstruct the No-Onset condition for proto-Berber too. Geminate-initial roots were clearly already geminate-initial in late proto-Berber (although Prasse, probably correctly, reconstructs them as *w-initial for pre-proto-Berber.) However, vowel-initial roots come from at least two sources: roots with vowel length (pre-proto-Berber h?) and roots with a glottal stop. The distinction is preserved in Zenaga, and t- shows up there in both cases. And, as it happens, Zenaga only allows the glottal stop in coda position. So it seems probable that late proto-Berber too allowed the glottal stop only in coda position.

Saturday, April 02, 2011

In search of the missing radical: a piece of Berber historical morphology

Berber normally has no glottal stops (ء = ʔ) – in fact, Chafik suggested that this was why North Africa favours the Warsh reading of the Qur'an, in which most glottal stops are omitted. However, it turns out* proto-Berber did have glottal stops - and you can still see their footprints on the verbal system.

Berber languages normally have three basic aspect/mood forms:
  • the “aorist” (or “simple imperfect”), used mainly for hypothetical events (“eat!”, “I will eat”, “I would eat”...);
  • the “preterite” (or “simple perfect”), used mainly for past events conceived of as wholes (“I ate”, “I have eaten”);
  • the “intensive” (or “intensive imperfect”), used for events ongoing at the time being referred to, irrespective of tense (“I eat”, “I am eating”, “I was eating”, “keep eating!”)
Usually, you can predict the preterite and intensive from the aorist. For three-consonant roots – eg lmd “learn”, a widespread Phoenician loanword – this is how it works in Tuareg (Tahaggart):
  • Aorist: ǎlməd “learn!”
  • Preterite: (y)-əlmǎd “(he) learned” (change the vowel pattern)
  • Intensive: (i-)lammǎd “he is learning” (double the middle consonant)
Tuareg has kept a distinction between two short vowels, ǎ and ə; but most varieties have just merged the two, so there is no difference in three-consonant roots between the aorist and preterite. So in Siwi, for example, you get:
  • Aorist: əlməd “learn!”
  • Preterite: (y)-əlməd “(he) learned”
  • Intensive: (i)-ləmməd “he is learning”
(Students of Akkadian/Assyrian/Babylonian will be getting a sense of déjà vu now...)

But some verbs have two consonants rather than three. Looking at Siwi I noticed that, if the verb had two consonants and no long vowels, there seemed to be two possibilities for the intensive, not just one; contrast:
  • Aorist: fəl “leave!”
  • Preterite: (y)-əfla “(he) left”
  • Intensive: (i)-təffal “he is leaving”
vs.
  • Aorist: ləs “wear!”
  • Preterite: (y)-əlsa “(he) wore”
  • Intensive: (i)-ləss “he is wearing”
So why the split?

Well, looking at the intensive forms, you see that in fəl you double the first consonant, while for ləs you double the second one. If you wanted to try to relate these to three-consonant verbs, you might think of something like:
- fəl < *Xfl
- ləs < *lsX

But if you look at Siwi on its own, there seem to be a lot of problems with this idea: in particular, why would the preterite of fəl end in -a?

Looking wider provides some answers. It turns out that in Tuareg – like Kabyle, and Tashelhiyt, and Ghadamsi, and a few other varieties – these verbs are distinct in the preterite too, and they are distinguished in exactly the way you'd expect from that little piece of internal reconstruction:
  • Aorist: əfəl “leave!”; əǵən “kneel!”
  • Preterite: (y)-fǎl “(he) left”; (y)-ǵǎn “(it) knelt”
  • Intensive: (y)-ffal “he is leaving”; (y)-ǵǵan “it is kneeling”
vs.
  • Aorist: ǎls “wear!”; əsəl "hear!"
  • Preterite: (y)-lsa “(he) wore”; (y)-sla "he heard"
  • Intensive: (y)-lass “he is wearing”; (y)-sall "he is hearing"
It's just that in Siwi – and Mzabi, and Chaoui, and Tarifit, and all the other Zenati Berber languages – the preterites of these two verb classes are merged, so they both end in -a. So our internal reconstruction is looking good... but what consonant might have been lost?

Zenaga, the Berber language of Mauritania, gives us part of the answer. In Zenaga, they look like this:
  • Aorist: ägun “kneel!”
  • Preterite: (y)-ugän “(it) knelt”
  • Intensive: (y)-uggan / (yə)-ttugun “it is kneeling”
vs.
  • Aorist: ätyši “wear!”, ätyšaʔ-m “wear! (to a group)”
  • Preterite: (y)-ityša “(he) wore; ityšäʔ-n “they wore”
  • Intensive: (yi)-yässä “he is wearing”; yässäʔ-n “they are wearing”
Notice that glottal stop ʔ that shows up when you add a consonant. That isn't automatic in Zenaga: contrast y-ugrah “he heard”, ugrān “they heard”. So it looks as though the original conjugation of “wear” was something like:
  • Aorist: *ǎlsəʔ “wear!”
  • Preterite: *(y)-əlsǎʔ “(he) wore”
  • Intensive: *(yə)-lassǎʔ “he is wearing”
We can also see that the missing first consonant in verbs like fəl, if they had one, was not ʔ – as far as I know, no Berber language has preserved evidence of what it may have been. (The t showing up in Siwi is probably not original, but rather borrowed from the intensive of vowel-initial roots.)

But there's still a problem here: why is *-ǎʔ reflected differently in the intensive vs. the preterite? A full answer for that would require a look at reflexes of the glottal stop in general, not just in the verbal system. But in several Berber languages, in fact, it's reflected identically. Compare, from opposite ends of the Berber world:

Tashelhiyt (southern Morocco):
  • Aorist: ls “wear!”
  • Preterite: (i)-lsa “(he) wore”
  • Intensive: (i)-lssa “he is wearing”
Awjila (eastern Libya):
  • Aorist: əsəl “hear!”
  • Preterite: (yə)-sla “(he) heard”
  • Intensive: (i)-səlla “he is hearing”
Clearly, Tashelhiyt and Awjila are not likely to form a subgroup! So my tentative interpretation would be that the form with -a is regular, and the form without -a found in Siwi, and Tuareg, and Kabyle, and almost every other Berber language between southern Morocco and Awjila is analogical – the intensive is always formed from the aorist, and it must have felt wrong to have one that looks as though it's based on the preterite. I've been looking at the always problematic subgrouping of Berber lately, and this would have interesting implications for that – it would suggest that Kabyle is more closely related to Zenati than to Moroccan Atlas Berber, since they share this innovation. But in Berber a lot of innovations seem to have spread areally, so it's scarcely conclusive.

* (All but the last bit of this post is an introductory summary of work by Prasse, Kossmann, and Taine-Cheikh that I've recently been digesting. It offers an interesting small-scale parallel to the story of Saussure's laryngeals.)

Friday, April 01, 2011

Tunisian Berber and language shift

It is not that easy to find information on Tunisian Berber, so I was quite happy to come across this PhD thesis free online: Berber ethnicity and language shift in Tunisia, by Hamza Belgacem. The author, himself from Douiret, estimates that only about 60,000 Tunisians still speak Berber, and the number is dropping as their children grow up speaking Arabic. He calls the surviving varieties Douiri, Cheninnaoui, Djerbi and Matmati, and argues that they together form a single Tunisian Berber "dialect" on a par with Kabyle or Tashelhiyt. (However, he offers no opinion on whether the extinct variety of Sened belongs with the rest, and forms this opinion on the basis of comparison to Kabyle and Moroccan varieties, but not Tumzabt or Chaoui or other geographically closer varieties.) This Berber community of southern Tunisia represent the remnants of a mostly Arabised tribal confederation, the Ouerghemma, which controlled much of southern Tunisia and parts of what became northwestern Libya until the French conquest.

He paints an interesting picture of a small minority language under the impact of modernity. Traditionally, the language was preserved by a number of factors tying the community together and excluding outsiders. The women of each community would marry only within it - not just among the Ibadis, but within the Maliki villages as well (as formerly in Siwa.) Some testimonies suggest that land was not sold to outsiders (a claim I also heard about Berber-speaking villages around Bechar.) Such ties are being loosened by modernity, as people emigrate and marry out and as the national state has taken on a more active role in the community with compulsory education and mass media. On the other hand, modernity, in the form of international media, also exposes the young to pan-Berber, or at least pro-Berber, ideologies, counteracting the low value placed on it in the national context.

Berber, and more specifically village, identity seems to have been maintained, with emigrants to Tunis maintaining close ties with other emigrants from the same village. But in terms of language, the balance seems to have tipped against Berber throughout Tunisia: "Some children of five years old could not utter a coherent sentence in TuB... Hardly any Tunisian Berbers under 30 speak TuB fluently but they may be able to utter a few words or understand what is said in Berber... hardly anyone under 10 years of age uses or knows TuB except for a few words or expressions", although there reportedly remain "certain clans, where the whole population still speak TuB, including all the children." There are a couple of pithy quotes from interviewees expressing why this happened: "Our language is excellent but it does not put bread on the table", "Our children are reluctant to speak our language outside the home because the other children of Arabophones laugh at them." The author suggests that Berber may survive in Tunisia if attitudes towards Berber continue to grow more positive, but that strikes me as a bit optimistic given his observations - which adds to the urgency of producing a decent description of the language.

Wednesday, March 09, 2011

Linguistic diversity in Libya

The Interim Transitional National Council of Libya has a website up now, at which you can watch representatives of various towns declare their allegiance to the revolution and/or transitional government (and, in at least two cases, explicitly say they don't want foreign intervention.) These statements, as one might expect given the official context, are essentially in Standard Arabic with few dialectal features (although the numbers tend to be pronounced fairly dialectally.) But the first statement, from Nalut in the Nafusa mountains of the west, has a surprise at the end: it turns out to be bilingual, with a Nafusi Berber summary given at the end (from 1:29 on), opening with Azul fellaken Ilibiyen, "Greetings, Libyans." A nicely-balanced gesture, that - strongly reaffirming national unity by pledging allegiance to a government that currently isn't even geographically contiguous with it, while also implicitly saying, in the face of years of Qaddafi's nonsense: we have our own language as well as Arabic, and we think it's appropriate for addressing the nation, not just for talking to each other. That balance - neither suppression of minority identities for the sake of unity, nor self-absorbed pursuit of minority rights while ignoring oppression affecting the whole country - strikes me as a good omen for Libya's future, if only they manage to end this war fast enough.

A very large majority of Libyans have Arabic as their mother tongue - in fact, Western Eastern Libya was described by the colonial anthropologist Evans-Pritchard as the most Arab place on earth outside Arabia itself. However, the country also has a noteworthy Berber-speaking minority (about 5%, if you dare to trust Ethnologue; it's not as though anyone's ever counted them in the past several decades.) Most speakers are concentrated in the northwest, where they (traditionally, for once) call themselves Imazighen: the port of Zuwara, along with many towns of the Nafusa mountains, such as Yefren and Nalut. All of that region - Arabic-speaking towns as well as Berber-speaking ones - is currently reported to be free of Qaddafi; language, thankfully, does not appear to be acting as a dividing factor there. A quite distinctive Berber language is spoken in the desert oasis of Ghadames on the Algerian border. There is a Tuareg community in the southwest, around Ghat and Ubari. The isolated Berber-speaking communities of Awjila in the southeast and Sokna near the middle are shifting to Arabic (this process is almost complete in Sokna) - their languages are of extreme historical interest and are very inadequately documented. Other longstanding linguistic minorities (the Muslim Greeks of Sosa, the Teda of the far south, etc.) are much smaller, numbering in perhaps thousands each. But for decades, Libya has been practically terra incognita for descriptive linguistic research: even work on its Arabic dialects has been scarce, let alone on politically sensitive minority languages. When (inshallah) the Libyans establish a stable and free state, it would be well worth documenting its linguistic diversity, both for better interpreting North African history and for informing Libyan educational policy.

Wednesday, March 02, 2011

From hatred to singing in two easy steps

In Kabyle, the word for "sing" is šnu. No other Berber language is known to have a similar word for sing (see Nait-Zerrad, s.v. CN), and both the verbal noun and its plural are formed on an Arabic pattern (ššna, pl. ššnawi); so one is almost forced to look to Arabic for its origins. But ask the average Arabic-speaker in modern-day Algeria, and they'll tell you they've never heard any such word.

In Classical Arabic, there is a fairly rare verb šani'a شنئ, meaning "to hate", probably best-known from the third verse of Surat al-Kawthar: 'inna šāni'aka huwa l-'abtar "For he who hateth thee, he will be cut off (from Future Hope)". (Cognate words are found elsewhere in Semitic, for example Hebrew śānē', Syriac snā "hate".) This has barely survived in spoken Arabic, but (according to de Prémare) the causative šənnā is still used in Tangier (Morocco), meaning "to taunt someone by showing him something he wants that you won't give him."

Phonetically, šani'a is a perfect match for šnu (the glottal stop/hamza becomes y in colloquials, and Arabic final-y verbs normally end up in Kabyle as final-u, for reasons I won't go into) - but semantically, surely this is absurd?

So I would have thought, until, idly browsing through a glossary of the rather conservative Bedouin Arabic dialect of the Nefzaoua area in southern Tunisia (Boris 1951), I found the following entry:
شنى šnệ... inacc. yẹ́šni...; noms d'act. šänyân et šạ́ni: 1) "critiquer en vers, faire la satire"... 2) "détester".

شنى šnē... impf. yašnī...; verbal nouns šanyān and šany: 1) to criticise in verse, to satirise... 2) to hate
"Hate" to "criticise in verse" is a credible change, and so is "criticise in verse" to "sing". Suddenly, a connection that looked impossible becomes almost obvious.

In this case, as in many others, Kabyle has preserved an Arabic word that almost every Arabic dialect in North Africa has lost - but to make sense of the connection you have to look at a wide range of Arabic dialects, not just checking Classical Arabic and stopping there. The converse also applies: when looking into Berber loans into an Arabic dialect, it's not enough to look just at the Berber spoken next door. People move around, and words that were familiar in one generation may be forgotten in the next one.

Of course, if the Nefzaoua data weren't available, there's no way you could accept a comparison like this - and, if several thousand years had passed since the word was borrowed, instead of less than 1500, that intermediate step probably would not have survived. In other words, semantic change can rather easily erase connections beyond any reasonable hope of retrieval. This is one of the main difficulties in long-range historical linguistics - the further back you go, the more cases like this.

Sunday, February 27, 2011

Linguistic Survey of India recordings

The Digital South Asia Library at Chicago have just put online for the first time the gramophone recordings originally intended to supplement the Linguistic Survey of India, collected 1913-1929. Burma is also included. If you are interested in almost any South Asian language, this cannot be passed up: Gramophone Recordings from the Linguistic Survey of India. It brings back memories of my time at the Rosetta Project...

Thursday, February 24, 2011

Two poems of the Libyan Revolution

A poem from western Libya in honour of the new revolution - in Berber, I think the Zuwara dialect - that sums it up nicely:
Taẓiḍərt af akud
D asirm g timalt n agdud
D xa yəṛwa ala yəffud!

Patience for the time
And hope for the future of the people
And he who is thirsty shall drink his fill!
(Note some linguistically interesting features: the use of d "and" to link clauses rather than noun phrases is a calque of Arabic wa- - in other Berber languages d normally only links noun phrases; and the future prefix xa derives from a shortening of yə-xsa "he wants", just as English "will" comes from a full verb that meant "to want".)

Poking around on YouTube reveals a fair number of very angry Arab poets' responses to Qaddafi, some from as far afield as Kuwait, but it took some looking for me to find one in Libyan dialect (contrast it to Saif's speech yesterday); here it is, "Poem for the free men of Libya:
ينصر الله الشعب في كل أوطانه
ويسخط الظالم و جميع عوانه
...
يكفي سنين تحت الظلام حزانا
اليوم نسقوكم من كاس المرار اللي زمان سقانا
زال الظلام وعدى اليوم زمانا

yənṣəṛ əḷḷāh əššaʕb f kəll 'awṭānah
u yasxaṭ əđ̣đ̣āləm u žmīʕ ʕwānah
...
yəkfī snīn taħt əđ̣đ̣ḷām ħazānā
əlyōm nəsgūkam mən kās əlmṛāṛ əlli zmān səgānā
zāl əđ̣đ̣aḷām u ʕaddā lyōm zmānā

God grant the people victory in all their lands
And cursed be the oppressor and all his helping hands...
Enough years in the dark have we already suffered thus
Now we serve you the cup of gall that you used to serve us
The darkness now has ended and our time has come at last
(Linguistic notes: the 2nd person masculine plural [kʌm] (and 3mpl [hʌm]) are characteristic - they were one of the features that struck me most in the speech of Western Desert Bedouins. The [g] for Classical /q/ is of course a pan-Arab feature of Bedouin dialects. I took some minor liberties with the translation to get it to rhyme.)

Monday, February 21, 2011

Gaddafi Jr's speech

In his rather desperate speech today, Saif Al Islam Gaddafi opened with a sociolinguistically very interesting statement:

əlyōm saatakallam maʕākum... bidūn waraqa maktūba, 'aw xiṭāb maktūb. 'aw natakallam maʕakum bi... luɣa ħattā ʕarabiyya fuṣħa. əlyōm saatakallam maʕakum bilahža lībiyya. wa-sa'uxāṭibkum mubāšaratan, ka-fard min 'afrād hāða ššaʕb əllībi. wa-sa'akūn irtižāliyyan fī kalimatī. wa-ħattā l'afkār wa-nniqāṭ ɣeyr mujahhaza u-muʕadda musbaqan. liʔanna hāðā ħadīθ min alqalb wa-lʕaql.(YouTube - first minute; conspicuously dialectal bits bolded)


Today I will speak with you... without a written paper, or a written speech. (N)or even speak to you in the Classical (fuṣħā) Arabic language. Today I will speak with you in Libyan dialect, and address you directly, as an individual member of this Libyan people. And I will speak extempore. Even the ideas and the points are not prepared in advance. Because this is a speech from the heart and the mind.


Now the explicit association between dialect, extempore speech, and speaking as "one of us" is fairly obvious, if interesting. But the odd thing is that this paragraph, like the rest of the speech, isn't very dialectal at all; it seems far closer to Standard Arabic than to any dialect. Some dialectal features are present, but a lot of unambiguously Classical constructions are used; even something as basic as the first person singular oscillates between Libyan n- and Classical 'a-. What it looks more like is some sort of intermediate ground between dialect and standard - or, if you prefer, like the highest level of Arabic that he is capable of extemporising in at short notice.

Readers may recall that Ben Ali tried the same gambit in his last speech (though Mubarak never resorted to it.) An omen? Let's hope so.

Monday, February 14, 2011

What it's like learning Darja

What little spare time I have left over these days is mostly dedicated to figuring out the fantastic things going on in the Arab world. Two months ago I would have said it was impossible that two dictators could be brought down by peaceful popular uprisings in such a short time - now anything seems possible. Siwa, by the way, is fine - they seem to have remained quiet the whole time under their shaykhs' cautious leadership (although an oasis nearer the Nile Valley, Kharga, suffered brutally when they tried to march.) So don't expect too many postings unless I come up with a new linguistic angle on the political situation...

However, one thing that's not changing in the Arab world is diglossia - so, to tide you over, here's a nice personal account of Moroccans' seemingly schizophrenic attitudes towards their own language that I came across the other day: Back in the Day.... Most of it carries over seamlessly to Algeria.

Tuesday, January 18, 2011

Language use in Tunisian politics

Unless you've been stuck on an iceberg in the Antarctic, you probably know that the Tunisian people have earned themselves imperishable honour, no matter what happens next, by kicking out their thieving, torturing control freak of an ex-president Ben Ali. Mark Liberman (via LH) has already commented on his unusual choice of dialect in his last speech. Fortunately, he's yesterday's news, so I'm going to comment instead on the language being used by the newly significant figures jockeying for power. Due warning: the sociolinguistics of politics is not my specialty, and I don't have much prior experience of specifically Tunisian language use, so read on at your peril and feel free to correct me if you have a better idea. For non-Arabic speakers, the key point to remember is that in any one country Arabic has at least two basic levels - formal Fusha and dialectal Darja - which are different enough grammatically and lexically to be considered separate languages, but which can be combined in appropriate circumstances.

The Prime Minister is Mohamed Ghannouchi. He first came to prominence on Saturday when he briefly declared himself acting President. This speech was entirely in Fusha - no efforts to add a personal touch here, simply officialese. The only dialectal features I notice are the pronunciation of jīm as ž, and of some short low vowels as ə. The delivery, however, is notably non-fluent - he's reading it slowly from a paper, pausing sometimes every three or four words, and he makes a mistake in case marking ('ad`ū kāffati 'abnā'i tūnəs "I call upon all the sons of Tunisia" - should have been kāffata.) Today, as Prime Minister he announced the new cabinet; his speech is a bit less halting (although still halting enough that you get several elision failures, like li al-ħayāti l`āmmah for lilħayāti l`āmmah), but as before it is entirely in Fusha and is being read out from a paper. The names, however, are pronounced in Darja, as they would be in conversation. Reminiscent of Chadli Bendjedid, this looks like the delivery of a politician who feels the need to speak Fusha for symbolic reasons but isn't actually fluent enough in it to do so impromptu - he was born in 1941, when Tunisia's educational system still operated largely in French. More tellingly, his delivery betrays the fact that he has never had the need to master rhetoric or appeal to a mass audience.

Moncef Marzouki, a secular leftist opposition figure calling for the old ruling party to get out, similarly sticks to Fusha throughout a recent interview with Aljazeera, avoiding dialect forms with remarkable persistence. His language use nonetheless contrasts strikingly with Mr. Ghannouchi's: Mr. Marzouki speaks quickly and fluently off the cuff, without consulting any visible notes, and without any conspicuous errors in delivery. Yet Mr. Marzouki is only 4 years younger than Mr. Ghannouchi, and, having studied medicine, undoubtedly did his university in French; has he simply been more motivated to learn to speak to a wide audience? The choice of consistent Fusha seems to reflect Aljazeera's pan-Arab audience; in an older video, aimed more at a Tunisian audience, he again speaks primarily in Fusha, but makes a number of shifts into Darja, for example evoking immediate reactions (eg, with Darja underlined: lākin anā lammā wužəht bihād əṭṭalab qult: āš nənžəm nḍīf 'anā? "But me, when I was faced with this request, I thought: "What can I add?") or quoting proverbs (eg sāl əlmužaṛṛab ma tsālš əṭṭbīb "Ask a person with experience, not a doctor") The effect, to me, is reminiscent of a classroom lecture.

The regime's favourite bogeyman for many years, the Islamist leader Rachid El Ghannouchi, has announced plans to return shortly, though not to run for office. In his speech of 2 days ago, he uses Fusha consistently and fluently, with an intonation reminiscent of a sermon, and shows only sporadic dialectal phonetic features (eg qámə` for qam` "repression"). Yet he shifts into Darja briefly (at about 4:50): after warning security forces that those who kill innocents will be damned to Hell, in the maximally formal language of a quotation from the Qur'an (wa-may͂ yaqtul mu'minan muta`ammidan, fa-žazā'uhu žahannamu xālidan fīhā, wa-ġaḍiba ḷḷāhu `alayhi wa-la`anahu wa-'a`adda lahu `ađāban 'alīmā "Whoso slayeth a believer of set purpose, his reward is hell for ever. Allah is wroth against him and He hath cursed him and prepared for him an awful doom"*), he suddenly caps it with a brief colloquial appeal to their common sense: əṭṭāġiya muš məš isədd a`līk "the tyrant isn't gonna save you". I can't hear any obvious traces of his southern origin (no g replacing q, for example), but I don't know Tunisian dialects well enough to spot subtler indications.

As for the protesters? Well, listen for yourself to one of the latest. Some slogans are definitely dialectal: Tūnəs, Tūnəs, ħəṛṛa ħəṛṛa, wa-t-tažammu` `ala baṛṛa "Tunisia free, RCD out!" Others are purely Fusha (though minus inconvenient case endings, as is common in less formal Fusha): yā tažammu` yā žabān, ša`b tūnəs lā yuhān "RCD you cowards: The people of Tunisia will not be belittled!"** Not hearing anything in French though, which is interesting given its prominent position in the Tunisian sociolinguistic environment: I suspect French would (rightly) be viewed as inappropriate for an appeal to the people of the nation, no matter how many people may speak it as a second language, whereas Fusha or Darja are equally suitable for demonstrations.

*: Stupid mistake corrected, and Pickthal translation of 4:93 substituted. It was getting late when I wrote that.
**: Looks like I misheard this one too! Corrected following Bilel's comments below. I guess transcribing YouTube videos is a risky business.

Saturday, January 08, 2011

Berber words in Roman times, and Ghomara Berber material

A couple of goodies for readers interested in North Africa / contact / the classical Mediterranean (if you fall into the first category, incidentally, you should also be following the major recent events in Algeria and Tunisia.):

Jamal El Hannouche, having finished his MA at Leiden, has recently put up Ghomara Berber: A Brief Grammatical Survey and Arabic Influence in Ghomara Berber. These are important reading for Berber philologists: despite its location in northern Morocco near the Rif, Ghomara Berber is not at all closely related to Tarifit, and shows some unusual features such as a feminine plural in -an. (The name of nearby Tétouan thus represents Ghomara Tiṭṭiwan, not Tiṭṭawin as other Berber-speakers might assume.) However, they are of even greater interest for contact phenomena: Ghomara Berber is one of very few languages (along with Agia Varvara Romani) to borrow fully conjugated verbs, from Arabic in this case. The only previous work on Ghomara Berber was a brief article in 1929 (and the Ethnologue has for some time been spreading the misconception that it is extinct); this is the first grammatical sketch of the language.

Carles Múrcia has recently completed his PhD at Barcelona, and put it up online: La llengua amaziga a l’antiguitat a partir de les fonts gregues i llatines. I'm afraid it's in Catalan, but if you can read French or Spanish you shouldn't have much difficulty (although it would be nice if he had translated more of the Greek quotations.) So far I've read the parts about Egypt and Cyrenaica. For Egypt, he points out there is no linguistic evidence that the Lebu / Libyans or Meshwesh, or any of the other Western Desert tribes recorded before the Mazices of the Byzantine era, spoke Berber, nor even that Siwa spoke Berber before the Byzantine era. This fits with my own observations that Siwi is simply too much like Western Libyan Berber to be the survival of an ancient Berber language of the Western Desert - although the activists who urge Imazighen to date their calendar from the "Amazigh" conquest of Egypt by the Libyans may not be happy with this cautious conclusion! For Cyrenaica, on the other hand, he shows that a number of words recorded in classical sources have convincing Berber etymologies, suggesting that Awjila may represent the continuation of a very early Berber-speaking population.

Interestingly, the words with Berber etymologies generally lack the characteristic Berber nominal prefix a-/ta-, which must still have been a separable word at that stage. For example, one Berber root that brought back memories of the Sahara is gelela, recorded by Cassius Felix as "coloquintidis interioris carnis" - the flesh of the inside of the colocynth, a bitter melon that grows wild in the Sahara and is commonly fed to goats. This corresponds to modern Tuareg tagăllăt, and to Kwarandzyey tsigərrəts, both meaning "colocynth" - but in those forms, the feminine prefix ta- (or ti-) has been added.

Wednesday, December 22, 2010

No word for heLLo?

It's no great surprise to find words in another language that have no English equivalent, if what they refer to is an object that's unfamiliar to most English speakers. For example, it's scarcely surprising if English has no word for "dates that aren't quite ripe yet, but that already ooze honey if you bruise them" (Kwarandzyey azMamweg); only a very small number of English speakers are familiar with date maturation stages, whereas practically all Belbalis are. It's a bit more interesting when you find that a phenomenon equally common in both cultures can be described by a fixed word or phrase only in one of them. Here's a case in point that came up in my latest fieldwork.

One of the basic states of mind in Kwarandzyey (and among the few to be retained from Songhay) is being heLLo. Songhay cognates (from *hollo) mean "crazy, possessed", which in Kwarandzyey is bA; the Kwarandzyey meaning of heLLo is quite different. This word is used (usually with a smirk) of people acting happy (leaping around, singing, dancing, etc.) or showing inordinate confidence, with no thought for consequences or respectability - Har ndza ghar ana hell-a bA ddzunets ka, "as if he was the only person in the world". Being full, or intoxicated, helps make people heLLo, but isn't essential. A heLLo person is generally said not to praise his Lord (asbayHemd an mulana si), ie not to appreciate that the causes of his happiness are contingent. Arabic translations suggested include colloquial SameT (literally "bad-tasting", but as a mental state more like "inconsiderate" or "silly") and classical Taaghii (as in "Nay, but verily man is rebellious (yaTghaa) That he thinketh himself independent!"). Here's a nice example of people acting heLLo (apologies to football fans - the example I was looking for was South Africans celebrating in the streets after Mandela's release, video of which was described to me as showing people being heLLo, but I couldn't find it):



Obviously, the mental state is at least as present in English speaking cultures as in Tabelbala - in fact, it might be reasonable to say that regularly achieving heLLo-ness is an important and widely socially accepted goal for British youth. But is there a word or fixed phrase corresponding to the concept in English? If you can think of one, feel free to suggest it!

(PS: Pardon the transcription - my computer is broken, and I can't be bothered to do all the cut-and-pasting it would take to fix the diacritics.)

Friday, November 12, 2010

Back to the Sahara

In the near future I plan to do some further travel in Algeria in the Sahara, to study more Kwarandzyey of course but also other languages of the region. Any readers of this blog in the area that want to meet up, feel free to email me - let's see if we'll be in the same place... For obvious reasons, posting will continue to be sparse until I get back.

Wednesday, October 13, 2010

A note on Azer

In the unlikely event that you've heard of Azer, a northern dialect of Soninke formerly spoken in the now Arabic-speaking region of Tichit and Walata in southeastern Mauritania, you may well have formed the impression - as I did initially - that it was heavily influenced by Berber, like the Northern Songhay languages are. If you know anything about Berber, a look at Monteil's article on Azer is sufficient to dispel this idea. If you don't, then chapter 3 of Long's thesis on Northern Mande, which I just came across, clarifies the issue nicely. This rather highlights the Northern Songhay problem: if centuries of close contact with Berber left Azer so little changed, why is Northern Songhay so full of Berber words?

Wednesday, October 06, 2010

Reporting language "discovery"

Turning on the BBC yesterday, I was surprised to hear a descriptive linguistics story, about the "discovery" by linguists on the Enduring Voices project of a previously unknown Tibeto-Burman language called Koro in Arunachal Pradesh: Indian language is new to science. Insofar as one can judge from the report, it sounds like Koro is clearly distinct from its neighbours rather than being an ambiguous dialect-continuum case, so this should be interesting for comparative Tibeto-Burman.

What struck my attention most is that this made it into the news! There have been a couple of discoveries of new languages in Africa over the past decade or so - Baka, for example, and Tondi Songhay Kiini. And the belated realisation that Bangime is a clear isolate, rather than a dialect of "Dogon", actually reshapes our picture of West African linguistic history much more than finding any of these languages has. Where was the news coverage of these? Have media attitudes towards the newsworthiness of "new" languages changed? Is it because they're in Africa? Or did the linguists in question simply not issue any handy press releases? Publicity is a hassle, frankly, and no one wants to sound like they're playing Indiana Jones. But stories like these are a big part of what gets people interested in linguistics in the first place, and the general public who fund most linguistic work, whether through taxes or donations, need to know what they're getting for their money.

Wednesday, September 29, 2010

Small vocabularies, or lazy linguists?

In Guy Deutscher's new book The Language Glass (which I'll be reviewing on this blog sometime soon) he claims (p. 110) that "Linguists who have described languages of small illiterate societies estimate that the average size of their lexicons is between three thousand and five thousand words." This would be rather interesting, if verified - but this statement is not sourced at the back, and is in any case too vague (what counts as "small"?) to be relied on as it stands. Does anyone have any idea where he might have got this figure?

I haven't found his source, but Bonny Sands et al's paper "The Lexicon in Language Attrition: The Case of N|uu" gives a nice table of Khoisan dictionaries' sizes, ranging from 1,400 for N|uu to < 6,000 for Khwe and 24,500 for Khoekhoegowab. She prudently concludes "The correlation between linguist-hours in the field and lexicon size is so close that no conclusions about lexical attrition can be drawn" - the outlier, Khoekhoegowab, is not only the biggest of the lot (with over 250,000 speakers), but had its dictionary written by a team including a native speaker over the course of twenty years. Given that "2,000 - 5,000 word forms (in English) may cover 90-97% of the vocabulary used in spoken discourse (Adolphs & Schmitt 2004)", it is not surprising that it should take disproportionately long to move beyond the 5,000 word range. However, she also points out that "Gravelle (2001) reports finding only 2,300 dictionary entries in Meyah (Papuan) after 16 years of study", suggesting that some languages may simply have unusually small vocabularies. Along similar lines, Gertrud Schneider-Blum's talk Don’t waste words – some aspects of the Tima lexicon suggested that the Tima language of Kordofan had an unusually small number of nouns due to extensive polysemy and use of idioms (I can't remember any figures, nor indeed whether she gave any.)

I'd be interested to see other discussions of the issue of differences in lexicon size and explanations for them. My Kwarandzyey dictionary (in progress) so far stands at about 2000 words - it would be encouraging to think that I might already have done more than half the vocabulary, but I very much doubt it!

Wednesday, September 22, 2010

Kouriya

I finally got my hands on an article I had been looking for for a while about the "Kouriya" language of Gourara (around Timimoun, Algeria): Rachid Bouchemit, 1951. Le Kouriya du Gourara, Bulletin de Liaison Saharienne 5, p.46-47. While short, it's significantly more informative than the vague rumours to be found in other sources. "Kouriya", it turns out, was the general-purpose name given locally to any Black African language - "L'unité du terme cache la pluralité des idiomes: Haoussa, Bambra, Foullan, Mouchi, Songhai, Bornou, Boubou, Gouroungou, Minka, Sarnou, Nourma, Kanembou, Karkawi, etc...", in particular as spoken by ex-slaves in the region. Following the abolition of slavery, these languages, no longer reinforced by the arrival of new slaves, rapidly fell into disuse; the new generation learned Arabic and Taznatit instead. By 1951, the author could find only seven or eight speakers of a "Kouriya" in Timimoun, and only two of them spoke the same language, namely Bambara.

While the author leaves the etymology unexplained, I would add that the term "Kouriya", and the corresponding ethnonym kuri, probably derive from Songhay koyra "town, village", used to form the Songhays' own name for themselves, koyra-boro "townsman"; Songhay is, after all, the nearest major ethnic group in the Sahel to the Gourara region.

Monday, September 13, 2010

Arabic right-hemispheric WEIRDness

Recently Language Hat asked for informed reactions to a BBC report claiming that Reading Arabic 'hard for brain'. The papers under discussion are to be found at Eviatar's home page, in particular the 2009 paper "Language status and hemispheric involvement in reading: Evidence from trilingual Arabic speakers tested in Arabic, Hebrew, and English" but also clearly the 2004 paper "Orthography and the hemispheres: Visual and linguistic aspects of letter processing". Now I'm no psycholinguist, but obviously this story smells fishy, so I had a closer look.

At least one glaring mistake seems to be clearly the BBC's fault: it wrongly claims "When the Arabic readers saw similar letters with their right hemispheres, they answered randomly - they could not tell them apart at all." In fact, this seems to conflate two different experiments. Telling letters apart was the first task in the 2004 paper, and the Arabic readers' error rates for similar letters were only 8% (Table 6) - worse than with the left hemisphere, but not nearly so bad. The claim that "there is a specific RH deficit in reading Arabic, because that is the only condition (with bilateral presentation), where these native Arabic speakers responded at chance" comes from the 2009 paper - but the task referred to there was substantially more complicated. They were looking at words/nonwords, not letters; they were presented with two words, one for each hemisphere, one of which was underlined; and they had to decide whether the underlined "word" was a real word or not. Other issues are not so much wrong as stupid: talking as though students could choose which hemisphere to learn with, for example.

However, the BBC cannot be blamed for drawing excessively sweeping conclusions from this experiment. The authors themselves talk of their results as applicable to Arabic in general, which rather overstates the case. In both papers, the Arabic speakers were all also fluent speakers of Hebrew, which they had studied since second grade, and were living in a state where Hebrew is the dominant language. In the 2004 test, at least, they were also all undergraduates studying degrees taught in Hebrew. Obviously, this is a rather unusual situation for Arabic speakers! In particular, it is one where pragmatic (and status-related) motivations to study Hebrew, and opportunities to familiarise oneself with it, are likely to be much greater than for Arabic (especially given the big difference between spoken and written Arabic.) In some types of tests, these speakers's right hemispheres seem to read Hebrew more easily than Arabic. The authors take this to mean that there is a "specific difficulty of the RH with Arabic orthography". But, without further testing elsewhere, it can equally well be taken to reflect the sociolinguistic situation of Palestinian citizens of Israel. This is, in fact, a special case of a much wider problem: most psychology experiments focus on "WEIRD" populations (read the link - it's a concept very much worth remembering when you read the science news.)

Friday, September 10, 2010

Doctorate done

Eid Mubarak everyone! I am now Dr. Souag. (As of a couple of weeks ago, actually, but I've been doing other stuff instead of being online.) You can read my thesis online, for the moment: Grammatical Contact in the Sahara. My examiners were Prof. Jeffrey Heath and Dr Martin Orwin. Thanks once again to everyone in Tabelbala or Siwa that helped me learn their languages, and to my supervisors, teachers, friends, and family. I'm currently working out future plans, but rest assured that they include plenty more research.

Thursday, August 19, 2010

Linguistic purism in 19th century Libyan Berber

Looking through Richardson's (1850) vocabulary of Sokna Berber today, I came across a wonderful little piece of sociolinguistic history. The vocabulary in question was written by a Sokni, Ali ben El-Haj Abd et-Tawil, with English translations added by Richardson. He wrote, among other things, the numerals. 1-3 are Berber (əjjin اجين, sən سن, šaṛəṭ شارط), while 4 is Arabic (أربعة arb`a). But when he reached 5 there was a moment of indecision:
Do you see what's going on there? He started out by writing خمسة xəmsa, the Arabic loanword meaning "five" - which, if other languages of the region are any guide, was the usual word for "five" in everyday Sokni. But then he had a thought - xəmsa is just Arabic, it's not proper Sokni, and I ought to be giving this stranger proper Sokni - and he overwrote the word with فوس fus "hand", used by Berber and Songhay groups through much of the Sahara (eg Siwi fus=hand, Kwarandzyey kəmbi=hand) as a substitute for "five" to prevent Arabic speakers from understanding, as they would if the normal numerals, borrowed from Arabic, were used. What at first sight looks like just a piece of messy handwriting turns out to bear witness to a moment of linguistic purism.

Saturday, July 03, 2010

The unreliability of Afroasiatic etymologies

The fact that Semitic, Egyptian, Berber, Cushitic, and Chadic all belong to a single family - Afroasiatic - is fairly secure, based on striking correspondences in basic morphology. However, it is often not appreciated just how difficult it is to find reliable lexical comparisons between these families, and just how primitive the current state of AA reconstruction is. The easiest source of AA etymologies online is Militarev's database on Starling, so I'm going to pick on it for this post (Orel & Stolbova and Ehret reveal similar issues, but the latter doesn't even include Berber, and I'm focusing mainly on Berber entries here for convenience.)

Suspiciously many entries are listed as having a cognate in only one Berber language (eg earth, hide, skin, run away); given the general closeness of different Berber varieties, you would expect valid proto-Berber terms to be reflected in more than one place. However, these could always be right. Other issues are more serious.

In several cases, a single proto-Berber root is split across several AA ones, due to mistaken sound correspondences. For example:
  • Proto-Berber *i-qăs "bone, (fruit) pit" is split between PAA *ʔayš/ʔawš- "ripened grain, corn" with Zenaga iʔssi (quoted without the glottal stop) "os; grain, graine, baie; comprimé, pilule, cachet, pastille; perle" (Taine-Cheikh), and *ḳ(ʷ)as "bone", with all other reflexes of *iqăs, even though Berber γ (<*q) commonly corresponds to Zenaga ʔ.
  • Proto-Berber *ta-Hăli (> *ti-Həli) "sheep" is split between pAA *ʔayl "ram" and *bawil "ram", although Ghadames-Awjila v corresponds regularly to Tuareg h and other Berber Ø. (A couple of forms, like Figuig tili mistakenly glossed as "ram", have even somehow found their way into a third etymon, "proto-Berber" *laH!) The issue is alluded to in a cryptic comment under the Berber section of PAA *waʔil "wild goat/ram; antelope": "Pr. H No. 220 (and Kössm. 193): Ghdm., Audj. Hgr etc. te-hele < *tiHeli, which, on the contrary, is to be connected with *ʔayl- 'ram' 3061 (together with Brb. forms of the t-ili type), as *ʔ > h in Hgr, while *ʕ > Hgr 0".

  • Most reflexes of pan-Berber ikərri / akrar "ram" are assigned to PAA *kar(w)- "ram, goat; lamb; kid". (The Semitic parallels listed for this word are rather interesting.) But Zenaga ǝgrǝrh, pl. gurănh 'bélier' (Nic. 156), on its own, is given a supposed proto-Berber form *gur- "ram", corresponding to an AA form *(ʔa-)gʷar "kind of antelope; ram; goat". In fact, however, there is a common correspondence of Zenaga g followed by a sonorant to proto-Berber k (eg ägärgur "chest" = Siwi ikərkər, əməgyih "dine" = Kabyle iməkli etc), and this word is obviously related to the other Berber forms.
Another case is listed as doubtful, eg:
  • Most reflexes of Proto-Berber *a-lăqŭm "camel" are under PAA *ʕalVḳ/g- ˜ *lVḳ/gum- ˜ *ḳalVm- "camel"; but the Zenaga one äyiʔm, with regular *l > y (in his source's transcription ǯ) and common *γ > ʔ as seen previously, ends up as PAA *gam-al- (?).
Similarly, unrelated forms may be grouped together due to accidental similarity, eg:
  • Under PAA *kʷay(-t)- "hen; partridge; dove; chick" is listed a "proto-Berber" form *i-kaHi; but the Ahaggar form listed corresponds regularly to Niger Tuareg tekažit, Mali Tuareg tekazzit, Awjila təkažit "hen" (see Kossmann 2005:60), and as such is unrelated to the Ayr and Tawllemmet forms takəyya quoted.
Another problem is undetected loans; this applies especially in sub-Saharan Africa, where little work has been done on their impact. PAA *ʔa/iw / *waʔ "bull, cow" is supported by Tawellemmet hawu "cow", isolated in Berber and obviously borrowed from Songhay, cp. Zarma haw, Tadaksahak hawú; removing this from the etymology leaves only pan-Tuareg iwan "cows", with no evidence for the desired *H. PAA *bar "cereal, corn" is supported by Zenaga būru "bread"; but this word is isolated in Berber and widespread in West Africa (eg Wolof mbuuru, Soninke buuru, Bambara nbuuru, Peul mbuuru, Zarma buuru), and is more likely a loan from Wolof or Pulaar.

Interestingly, most of the problem cases I've noticed in this quick skim are related to agricultural terminology. I wonder if that has anything to do with the particular interest of such terms for archeologists motivating a more intense search for cognates.

Thursday, June 24, 2010

Why they thought the Berbers came from Yemen

A long-standing tradition in North Africa, convincingly rejected by Ibn Khaldūn but perpetuated by poets and curricula alike, claims that some major Berber tribes descend from Yemeni Arabs through semi-mythical pre-Islamic kings and their wholly mythical vast conquests. This idea has little to support it, and probably became popular because it allowed these tribes to claim prestigious connections in the context of a high culture dominated by Arab ideas; but why should the connection be specifically Yemeni, rather than, say, North Arabian or perhaps Persian? Linguistics suggests a possible answer.

In southern Arabia live several groups, most famously the Mehri tribe, whose languages, though Semitic, are only distantly related to Arabic, and quite incomprehensible to other Arabs. (You can hear recordings of it at SemArch.) Recently I borrowed a copy of the recently published Mehri Language of Oman, by Aaron Rubin; looking through it, I could see several points where Mehri resembles Berber but not Arabic that a traveller might seize on, notably:
  • -s ـس "her", -sən ـسن "their (f.)"; compare Siwi -nn-əs ـنّس "his/her", -n-sən ـنسن "their (m/f)". A 3rd person in -s was found in proto-Semitic, as shown by Akkadian, but was replaced in Arabic.
  • əl ال "not" (preverbal first element of negative); compare Tumzabt ul أُل. Again, this is found in Akkadian and hence must be proto-Semitic.
  • -ət ـت feminine singular; compare Siwi -ət ـت (feminine singular in Arabic borrowings.) Again, the connection is real, but dates back to proto-Semitic rather than indicating any special relationship between the two.
  • -tən ـتن feminine plural; compare Berber -tən ـتن (plural of some masculine nouns)
  • a- أَ used as a definite article for some nouns; compare Berber a- أَ(masculine singular noun prefix). A striking case is Mehri a-məsge:d أَمسجيد vs. Siwi a-məzdəg أمزدج "the mosque". However, in Mehri this indicates definiteness, and does not depend on gender; this is probably a coincidence.
  • tə-...-əm تـ...ـم second person plural imperfective, eg təkə́tbəm تكتبم "you (pl.) write"; compare Berber t-...-m تـ...ـم. The t- is cognate; not sure about the history of the -m offhand.
  • 'ār آر "except, but"; compare Tuareg ar.
  • ā آ "oh" (vocative); compare pan-Berber a أ. (This is actually found in Classical Arabic as well, أ, but is not widely used.)
None of these similarities in fact imply any close relationship between Berber and Mehri, of course; some are coincidental, while others can be traced back to proto-Semitic, and hence constitute evidence connecting Berber with Semitic, not specifically with Mehri. However, a medieval traveller between Yemen and North Africa would not have known that, and could easily have observed similarities like these and leapt to the seemingly plausible conclusion that Berber was connected to the language of these Yemeni tribes, who, like many Berbers, seemed to live just like Arabs yet speak totally differently.

Tuesday, June 15, 2010

The Berber language of Sokna (Libya)

Thank you SOAS library - I finally got a copy of Il dialetto berbero di Sokna! Sokna (they even have a Facebook group) is a small oasis south of Sirt in Libya, whose dialect of Berber, along with that of nearby El-Fogaha, is Siwi's closest relative. There were several surprises inside, including unusual vocabulary like amerru "mountain" or imeγri "Dhuhr (the midday prayer)", and some striking features shared with Siwi; one of the main ones is an unexpected bit of allomorphy. Across Berber, the second person plural ("you guys") is expressed on the verb with t-...-m, except in the imperative; Sokna does the same, so for example "you have" is t-la-m. In the imperative, you have a suffix -t; Sokna again does the same, eg sag-it-ten iyi-leḥbes "(you guys,) take them to prison!" But if you add an indirect object pronoun ("to him" etc.) to the imperative, you replace this t with an m, like the m in the second half of the non-imperative forms: eḍbeḥ-im-as a-na-dd y-used "(you guys) tell him to come to us!" The same thing happens in Siwi, except that in Siwi the prefixed t- of the non-imperative forms has disappeared. I'm doing a paper on the development of indirect object agreement in Siwi for the Berberologie conference in July, and this is a useful pointer to its history. Amazigh readers - have you come across anything like this?

Sadly, Berber is probably no longer spoken in Sokna. When this article was written in 1911, the shaykh of the oasis reported that only 4 or 5 Isuknan could still speak it, although many more could understand a bit. I don't know whether the people of Sokna today regret the loss of their language or are glad of it - but its disappearance destroys a key not just to Sokna's history but to that of Libya, Egypt, and the whole of North Africa, leaving only this article's fairly short wordlist (and a few even shorter older sources) as evidence for migrations between central Libya and Siwa and early contact with vanished pre-Sulaymi Arabic dialects.

Wednesday, June 09, 2010

Religious origins of the "Welsh Not"?

A well-known weapon in the arsenal deployed by educational systems the world over against local languages was what in the UK used to be called the Welsh Not - a piece of wood hung around the neck of a student caught speaking their own language, and passed on through the day to anyone that student heard speaking their language, so that whoever was wearing it at the end of the day would be punished. At a talk yesterday I heard that the same idea was implemented in Japan (against Ryukyuan languages) and Sudan (against Nubian.) Coincidentally, I just came across an account that gives interesting insight into the origins of this oppressive practice:
"With a general consent of all our company, it was ordained that there should be a palmer or ferula which should be in the keeping of him who was taken with an oath; and that he who had the palmer should give to every one that he took swearing, a palmada with it and the ferula; and whosoever at the time of evening or morning prayer was found to have the palmer, should have three blows given him by the captain or the master; and that he should still be bound to free himself by taking another, or else to run in danger of continuing the penalty, which, being executed a few days, reformed the vice, so that in three days together was not one oath heard to be sworn."The Observations of Sir Richard Hawkins, Knt in his voyage into the South Sea in the year 1593
Hard to imagine a ship full of sailors submitting to such a practice! But was this the original purpose of the Welsh Not? It would be interesting to find out. If anyone has an older citation to compare, I'd love to see it.

Thursday, May 20, 2010

Endangered languages on Aljazeera

Aljazeera English is doing an interesting series on language endangerment and revitalisation:
* Language on the brink, talking with the last speaker of Wichita.
* Saving the language of the Cherokee, in Tahlequah
* French region aims to save language, on Breton
* Turkey's fading linguistic heritage and Saving Turkey's Laz language, on Laz (a close relative of Georgian, not "an ancient tongue that bears no resemblance to any other language in the region".)
* Circassians in bid to save language in Jordan - at a talk this week by Enam al-Wer I heard that, at the start of the twentieth century, the only permanent population in Amman was Circassian.

Thursday, April 29, 2010

Manatees and bilingual compounds

In Djenné Chiini, the Western Songhay dialect of Djenné in Mali, the word for "manatee" is ayuumaa. This is clearly a compound of two elements: ayuu, the word for manatee throughout the rest of Songhay (as well as in Hausa), and maa from Bozo máa, which also means "manatee" (Bozo being the original language of the Djenné region.) It's as if the American English word for an elk were "elk-moose". I can't think of any other examples of this kind of half-borrowing, where a native word is "expanded" by adding on its translation into another language; can you?

(Sources: Daget 1953, La langue bozo; Heath 1998, Dictionnaire songhay-anglais-français, tome II: Djenné chiini.)

Monday, April 05, 2010

More on the WOLD Kanuri entry

The World Loanword Database is a great resource, and the Hausa/Kanuri team deserve congratulations for undertaking the Herculean labour of putting together two sets of etymologies. However, there are some issues with the Arabic etymologies in the Kanuri entry. The transcription is inconsistent and sometimes incorrect; more seriously, a few entries give incorrect meanings or impossible etymologies, as in the following cases:

3.592 àkú parrot: the quoted Arabic form is almost impossible as a Classical Arabic noun (and not in the Lisan al-Arab; the Arabic word is babγā’), and parrots are known in the Arab world only as an exotic import. Assuming the form exists in some Arabic dialect, it must be a loan from a sub-Saharan African language, not vice versa.
9.24 mágàsù scissors: the g and the u both suggest that this word entered directly from (Bedouin) Arabic, not via Hausa.
11.12 hàláltə́ own: if this is correctly transcribed, surely it comes from Arabic ħalāl “licit; one’s lawful property”. Arabic halak means “perish”.
11.79 ríwà dìò to earn: “ribā” means usury, and is strongly condemned in Islam; it is unlikely that this would be adopted as a neutral word “earn”. The more plausible source for both the Kanuri and the Hausa is Arabic ribħ “profit, gain”.
11.78 àlwúsùr wages: Perhaps < Arabic al-`ušr "tithe (< one-tenth)"; surely not from ma`āš.
14.451/6 kàjílí evening: “kajir” is not a possible native Classical Arabic word, and is not attested in Classical Arabic. If it’s in Shuwa, it must come from Kanuri, not vice versa.
16.34 tə́wə́rítə́ regret: Hausa tuubaa does come from Arabic, but clearly from Arabic tūb “repent”; it has nothing to do with Arabic ta’assaf (not *tāssaf) “regret”.
16.69 gàfə̀rtə́ forgive: the connection to Arabic γafar- is obviously correct, but Arabic yaʕfū is equally obviously not relevant; even if ʕ were normally reflected as g in Kanuri, it would leave the r unexplained.
18.33 kàsàttə́/àrdìtə́ admit: the Arabic form “kasat” does not exist. yarḍā means “may He hope/ approve” (as noted), not “admit”, making the connection rather tenuous.
18.45 áwúlò dìò boast: there is no Classical Arabic word “awulo”.
19.47 àmàrtə́ permit: Arabic ʔamar- means “he ordered”, not “permission”.
20.31 súlwé armor: Arabic silāħ means “weapons”, not “armor”.
21.24 àlàptà swear < ħalaf "swear" (not < allāh "the god")
21.37 àzáwù punishment: from Arabic ʕađāb “punishment, torment” rather than jazā’.
21.47 perjury: by what chain of semantic changes could “perjury” derive from “lawful”? And why would l > k?

Probable Arabic loanwords not listed as such include:
11.54 bàyîl stingy: from Arabic baxīl.
4.89 sûm poison: surely from Arabic samm?
4.93 sə̀lé bald: surely from Arabic ‘aṣla`?
5.26 kóló pot: perhaps cp. Arabic qullah (or onomatopeic?)
7.58 kábbì arch: surely from Arabic qubbah?
14.25 bàdìtə́ begin: surely from Arabic bada’?
11.29 lòrùtə́ damage: from Arabic ḍarr (impf. -ḍurr-). Cp. “judge” for ḍ > l.
24.02 wàltà become: perhaps from Maghrebi Arabic wəlli “become, return”.

In some cases, looking more widely allows the etymologies to be improved:
3.11 lə̀mân animal: < al-māl- "livestock, money", rather than al-mann "favor, benefit". For the dissimilation, compare the common Maghrebi Arabic change of n...n to n...l, eg badənjal < bāđinjān, fənjal < finjān.
2.34 lòrúsà wedding: probably from al-`arūs “bride” (Maghrebi Arabic l-aʕṛuṣa), rather than direct from ʕurs. Cp. Siwi aʕṛus “wedding”, with the same semantic shift.

There are also a few cases, many probably originally formatting issues, where the correct form is given in comments, but contradicted elsewhere:

3.25 sheep: the source cited, Kossmann 2005 (67), points out that the form quoted by Skinner, *adaman, is unattested. The correct form, adəmman, is found in Arabic as well as Berber, and refers to a type of sheep said to come from sub-Saharan Africa. Given that it refers to a specifically sub-Saharan sheep breed, 5 would seem a better classification than 4, though 4 is understandable.
3.78 camel: Kossmann 2005, cited, makes it rather clear than an Arabic origin for this word is very improbable. Moreover, there is no such Arabic word as “ləγəmal”; only the form jamal is correct.
4.87 physician: If Shuwa Arabic or some such variety has a term liktaay, there can be little doubt that it is a loan into Shuwa, not from Shuwa. As the comment indicates, this comes from English, not from Arabic.
7.422 blanket: The comments indicate a Berber form abroγ, but the field gives abrok. The Arabic etymology is less implausible than it appears, since the semantic shift to “full body covering” is well-attested, as in English “burka” from the same source.
12.081 above: here it is called areal and probably not Arabic, but under “sky” and “heaven” the same word is listed as “clearly borrowed”. One of these statements must be wrong.
13 zero: the Hausa form is transcribed correctly in comments, but wrongly under “Source words”.
18.51 write: rubuta is Hausa, not Berber, as the sources quoted make clear. The proto-Berber form had no suffix -t (as Kossmann indicates), and neither do any of the equivalent modern Berber verbs.
19.62/20.11 quarrel: If it’s related to “alhilaafu”, the Arabic form is al-xilāf. If it’s related to “judge”, that form is irrelevant. In either case, there is no Arabic word “alwalaʔ” with appropriate meaning.

Monday, March 01, 2010

Identify the language of this manuscript

A scan of much of the manuscript MS Leiden Or. 14.052 is available online. The main text of this manuscript is in a rather poor Arabic. The marginal and interlinear notes, however, are "in one or more West African languages", as yet unidentified. My best guess is that they're in Mandinka, based on the orthography's use of tanwīn and on the frequent word-initial a/i (suggestive of Mande's 3rd person subject pronouns), but I'm not sure; I haven't been able to decipher any phrases. Anyone else feel like having a look?

Tuesday, February 16, 2010

Subjacency: The judgements

Thank you very much for your responses, everybody! (If you haven't answered yet and want to, please do it before reading the rest of this post.)

Chomsky's intuitions were as follows (* marks ungrammaticality as usual):
  1. * That's the boy who they intercepted John's message to.
  2. * That's the boy who he believed the claim that John tricked.
  3. * That was a lecture that for him to understand was difficult.
  4. * Which book did John wonder why Bill had read?
  5. √ Which book did John think that Bill had read?
  6. √ What would you approve of John's drinking?
  7. * What would you approve of John's excessive drinking of?
Mine were that 1, 4, 5, 7, and (only after some thought) 6 were good, while 2 and 3 were wrong - but I exclude those judgements here, since I was reading the book and might have been swayed by my reactions to the arguments. My sister found 1, 2, and 4 wrong, 3 "weird but comprehensible", and 5-7 good - so even within a single family judgements vary significantly. Your 11 collective judgements (plus some friends and family, and excluding non-native speakers) add up as follows (grading "uncertain" as 0.5):


The discrepancy, and the level of individual variation, are striking - not a single reader agrees with all of Chomsky's judgements, and the only consistent judgements are 2 (always wrong) and 5 (always right.) Most of Chomsky's judgements also happen to be predicted by his (and others in the generative tradition's) theories; your judgements therefore often pose problems for those. According to Chomsky, 1 and 2 should both be ungrammatical for the same reason - they involve movement past more than one "barrier" (boundary of a 'noun phrase' (DP) or clause excluding the complementiser (IP)) at a time. Yet more than half the people here (including me) accept 1, while nobody accepts 2; one could argue that 2 should be less acceptable than 1 because it crosses three barriers rather than two, but why should 1 be acceptable at all? 4 should be ungrammatical because "why" is occupying a position that "which book" should have to move through - but about half of you (including me) think it's fine. And most readers of this blog find 7 to be better than 6 - the opposite of Chomsky's judgements and of the predictions of the "A-over-A" principle he was working with then (although the latter is obsolete.)

Chomsky (1963:51) said of sentences like these: "In some unknown way, the speaker of English devises the principles of [wh-movement etc.] on the basis of data available to him; still more mysterious, however, is the fact that he knows under what formal conditions these principles are applicable... The sentences of [1-3] are as 'unfamiliar' as the vast majority of those that we encounter in daily life, yet we know intuitively, without instruction or awareness, how they are to be treated by the system of grammatical rules which we have mastered." This seems to be false; individually we often find it difficult to decide the grammaticality of sentences like these, and collectively we routinely disagree on them. Certainly it cannot be construed as belonging to that part of the "knowledge of language" that is, in the words of Chomsky (1963:64), "independent of intelligence and of wide variations in individual experience".

If it did, then that would be rather interesting: it has been claimed that the principles of Subjacency must be innate, because children aren't exposed to enough evidence to deduce them otherwise. But given the level of variation actually observed, it is tempting to reverse the reasoning: children don't deduce most of the principles of Subjacency, so they must neither be exposed to enough evidence for them nor have innate knowledge of them. Rather than postulating arbitrary rules hard-wired into the brain and specific to the language faculty, a more promising way to explain Subjacency phenomena might be to try to derive them from processing difficulties, as suggested by Sag et al.

Subjacency intuitions

I've been reading an old Chomsky book, Language and Mind, lately. As usual, the moment he starts discussing what would eventually be called subjacency I find my intuitions are systematically different from his, and I'm curious: how common is this? By way of testing, here's a few sentences in English: which ones would you consider ungrammatical/unacceptable as phrased?
  1. That's the boy who they intercepted John's message to.
  2. That's the boy who he believed the claim that John tricked.
  3. That was a lecture that for him to understand was difficult.
  4. Which book did John wonder why Bill had read?
  5. Which book did John think that Bill had read?
  6. What would you approve of John's drinking?
  7. What would you approve of John's excessive drinking of?
Chomsky's grammaticality judgements will be provided later - they're on pp. 50-54 of the book.

Thursday, February 11, 2010

Berber manuscripts in Arabic script online

A major collection of early Tashelhiyt manuscripts from the 16th century onwards has gone online: Manuscrits arabes et berbères du Fonds Roux. It includes a copy of al-Hilali's Berber-Arabic lexicon. The Lmuhub Ulaḥbib library of Bejaia has also put a number of works online, including an 18th/19th century manuscript on theology in Kabyle: العقيدة السنوسية. Both collections are also of interest for their many Arabic books, but the Berber ones are particularly significant due to the serious paucity of materials for the study of precolonial Berber writing traditions.

Friday, February 05, 2010

Word Loanword Database

I shouldn't really be blogging at this stage of my thesis-writing, but this I had to share: the World Loanword Database has come online. Vocabularies likely to be of particular interest include Tarifiyt, Hausa, Kanuri, Iraqw, but there are plenty more, all carefully analysed for loanwords... Have fun, and feel free to discuss any mistakes you think you spot in it here :)

(Via Glossographia.)

Wednesday, January 27, 2010

Language endangerment: thoughts from Igli

I recently found a forum for the town of Igli, about 150 km north of Tabelbala as the crow flies. Igli's traditional language is a Berber variety called "Tabeldit", or in Arabic "Shelha" شلحة, reasonably close to the better-documented dialect of Figuig across the border but with significant differences (such as the first person singular in -ɛ rather than -γ.) In Igli, it is at least as endangered as Kwarandzyey, and is likely to disappear in another couple of generations - although I was told that it is doing better in the small neighbouring town of Mazzer. I think the reason, as in Tabelbala, is that parents started speaking only Arabic to their kids in the hope of giving them a head start in school, but all I know about Igli I heard from Glaouis in other towns. In situations like this, speakers inevitably see their language's disappearance with mixed feelings, and the following pair of posts forms a microcosm of the global language preservation debate:
The "Xiṭ Azugar" Project (posted by Shayma)

"Tabeldit Shelha is part of the fragrance of the Saoura region... a treasure inherited from our ancestors. Shall we preserve it, or let it disappear before our eyes?.... A secret weapon that saved some of us from death. How long will we remain with our hands tied as our language disappears before our eyes? Until when, until when?

I hope that these words have awakened your sleeping hearts and moved your sentiments. Therefore I present to you today this project, consisting of the establishment of an "Arabic-Shelha" dictionary to preserve our language. Therefore I ask the director and administrators and even the members to study this project; if you accept the idea, then let's start to lay down precise plans to overcome difficulties... and if you don't accept the suggestion, then we will do our ancestors an injustice... I urge you to take the matter seriously. To the administration, and all the members, let us put hand in hand. No more lamentation over Shelha, that doesn't help. What helps is effective work.

Forgive me for my harsh words, and I hope you accept the idea. The project is called "Xiṭ azugar" for historical reasons, because these words have saved a person from certain death.
This suggestion was acclaimed and adopted, and there is now a small Arabic-Shelha Dictionary forum. However, there was also some scepticism - the following post started a vigorous debate:
What would we lose if Shelha becomes extinct? (posted by igliab)

Following the increased concern with the local dialect "Shelha" from the brother members, for which thanks are due, I decided to pose the following question: What would we lose if this dialect became extinct?

It's not a language of civilisation, nor a language of science. And supposing we are able to make an "Arabic-Shelha" dictionary and lay down the rules for this language, will our sons agree to learn it? What would the motive be? It's not used at home, nor in public places. Or do we want to put it in museums and say we have "saved" it?

Moreover, by my reckoning those who speak it today are:
90% old men - 8% middle-aged men - 1.5% youths - 0.5% children. Admittedly I haven't made a study to come up with these figures but it could be worse than I anticipate, so it can be said that Shelha has no future in Igli.

I also told myself that if everyone thought the way I think then they would put down their pens and wait for the demise of Shelha, the way an ill man who has despaired of his state waits for death. But I rethought the issue, this time positively, and realised the need to put together a plan for its preservation. But what is the point of solutions if there is no logical, powerful reason, so the first question we have to answer is: why should we preserve Shelha? I urge the brothers to think deeply about this issue and put sentiments aside.
What would your thoughts be? Have you had a parallel experience?

Monday, January 11, 2010

Ajami in Boston

The Boston Globe has an article today about Ajami, the tradition of transcribing African languages in the Arabic script. It focuses particularly on the efforts of Fallou Ngom, whose work has been mainly on Wolof Ajami in Senegal, the subject of one of my first posts here. In the article he emphasises the potential historical significance of such work in opening up neglected sources on African history. While most African manuscripts are in Arabic, some historically rather interesting Ajami sources are known; for Mandinka, published historical manuscripts include the Pakao Book and the Bijini manuscript, the latter outlining regional history over the past 500 years. There are undoubtedly more out there that have gone uninvestigated simply for lack of enough historians who can read them. My work on Ajami has focused more on issues of orthography, however: most African languages have rather different sound systems to Arabic, and it's quite interesting to see what kind of devices they developed to make the alphabet fit better.

Saturday, January 09, 2010

Earliest Kwarandzyey source online (also Tarifit of Arzew)

It turns out that the earliest and most extensive published source on Kwarandzyey (Korandje), the language of Tabelbala in southwestern Algeria which I am studying, is downloadable online:

* Cancel, Lt. 1908. "Etude sur le dialecte de Tabelbala". Revue Africaine 52.

Readers may also be interested in Biarnay's study of the probably extinct Tarifit dialect that was then spoken at Arzew, in volumes 54 and 55 of the same publication.

Saturday, January 02, 2010

Siwi Scarborough Fair

Over the dinner mentioned in the last post I was also shown a Siwi poem sent as a text message - it's a rather below average example of the genre, but interesting as an representative illustration of Siwis' orthographic preferences.
كان تازمرت تجبد تيني
كان تفكت تعمار تازيري
كان اتغت تيرو اغي
كان امان نلبحورا يسقلبن اخي
كان الغم ينسخط ايزي
بردو شك غوري (غالي)
Or in Latin Berber orthography:
Kan tazemmurt tejbed tayni
Kan tfukt teɛmaṛ taziri
Kan tγatt tiṛew aγi,
Kan aman n lebḥuṛa yesqelben axi,
Kan alγem yensxeṭ izi,
Beṛdu cek γuṛi "γali".


So I decided to render it into English, taking a few liberties to reproduce the rhyme (for added faithfulness, change "flea" to "fly", and eliminate "someday" and "or three"):

If dates can come from an olive tree,
If the sun someday a moon shall be,
If a goat gives birth to a calf or three,
If milk fills the waters of every sea,
If a camel can turn itself into a flea -
Then only will you be dear to me.

Thursday, December 31, 2009

Siwi and Kabyle: same language family, but not same language

Just back from a nice evening with the Siwi community of Qatar. A Kabyle friend came along (hello if you're reading this!), giving me a chance to see first-hand to what extent Siwi and Kabyle are mutually comprehensible. The answer is: very little indeed. Looking through basic vocabulary it's not hard to find cognates; but when it comes to even short sentences, mystified expressions on both sides were the order of the day. The Berber languages of Algeria and Morocco may shade into one another to some extent, even across sub-family boundaries - there seem to be dialects for which it is difficult to decide whether they should be called Kabyle or Chaoui, for example. But by the time you get to Siwa, it's quite clear that you're dealing with a different language, even by Arabic speakers' rather generous standards. Further confirmation, if any was needed, that Berber is a language family, not a language.

Saturday, November 21, 2009

Songhay and Nilo-Saharan

Following up on the preceding post, I've been looking at Greenberg's (1966) Nilo-Saharan comparisons - specifically, the 29 ones involving Songhay that have reflexes in Kwarandzyey, the Songhay language least likely to be involved in recent contact with Nilo-Saharan. Of these, 20 have comparanda in Saharan (Kanuri/Kanembu + Teda/Daza + Berti + Beria/Zaghawa), 17 in Eastern Sudanic (Nubian, Nilotic, Surmic, etc.), vs. a maximum of 13 for any other branch. (At least 7 also have plausible Mande comparisons.) Now, Saharan only consists of about 4 languages (9 by Ethnologue standards.) For Eastern Sudanic, excluding Kuliak, the Ethnologue counts 103 languages, and a huge amount of internal diversity. If Songhay were equally distant from the whole of Nilo-Saharan, you would expect far more cognates with Eastern Sudanic than with Saharan; the figures suggest that the link (whatever its nature) is primarily with Saharan, and only secondarily, if at all, with the rest of the languages he classified as Nilo-Saharan.

The grammatical comparisons that Greenberg offers are interesting but not compelling; there are only 10 of them (only 4 with Kwarandzyey reflexes), and they often incorporate misrepresentations (as Lacroix noted, for example, -ma forms verbal nouns, not relatives/adjectives, and 1sg ay < *agay, reducing the similarity to forms like Zaghawa ai.) Some of the lexical ones, however, are rather good; similarities such as Koyraboro Senni kokoši “scale (of fish)” = Manga Kanuri kàskàsí “scale (of fish)” cry out for explanation, and, though quite rare, look sufficiently numerous that chance seems unlikely. But whether they should be explained by contact or borrowing remains unclear. Either scenario would be historically interesting, since at present rather a large expanse of Tuareg and Hausa-speaking land separates Songhay from even Kanuri, and Saharan originated closer to modern-day Darfur than to Lake Chad.

Sunday, October 18, 2009

Arabic loanwords in "proto-Nilo-Saharan"

Ehret 2001 (or see Nostratic.ru) looks at first sight like an astonishingly detailed reconstruction of Nilo-Saharan, with nice binary splits and loads of technology-related words for archeologists and anthropologists to sink their teeth into. Why shouldn't specialists take advantage of this amazing opportunity to correlate historical developments to linguistic ones?

I just found a handy answer to that question. Bender (1997:175ff) gives the 15 cognate sets in Ehret 2001 that are represented in the most sub-families of Nilo-Saharan. 3 of the 15 look distinctly like Arabic loans.

1387 *wàs “to grow large”: Fur wassiye “wide” and Songhay wásà “to be wide” are both from Arabic wāsi`- واسع. The other items cited – Ik “stand”, Kanuri “yawn”, Kunama “increase, augment”, and Uduk “to tassel, of corn” – are scarcely obvious candidates for being related to one another in the first place.

1297 *là:l “to call out (to someone)”: Kanuri làn “to abuse, curse” and Songhay láalí “to curse” are obviously from Arabic la`an- لعن; Kunama lal- “to denigrate” might be from the same source. That only leaves Uduk “to persuade, incite to do something” and Proto-Central-Sudanic “to call out”.

718 *t̪íwm “to finish, complete”: almost certainly Songhay tímmè “to be finished”, very likely Uduk t̪ím “to finish”, Ocolo t̪um “to finish”, and maybe even Fur time “total”, are from Arabic tamm- تمّ (impf. -timm-), as Bender (ibid:177) considers probable. That leaves Proto-Central-Sudanic, Kunama, and Maba “all”, Kanuri “ideophone of dying animal” (!), and Proto-Kuliak “buttocks”. The “all” set looks rather promising – the whole etymology, not so much.

There are plenty of other Arabic loanwords in Ehret's “Proto-Nilo-Saharan” – a particularly egregious example is Kanuri zàmzàmíyɑ̀ “leather bottle-shaped water vessel for journeys” (#1223 *zɛ̀m “to become damp, moist”), and other especially clear-cut cases include #1173 < sawṭ, #1185 < šamm – but the fact that they include a significant proportion of the best cognate sets is what really strikes me. If a reconstruction attempt can't distinguish a widely distributed recent loan from a cognate set that split more than eleven thousand years ago, any information it gives about readily diffused items like technologies is completely unreliable. For another review from a similar perspective, try Blench 2000 (not sure why it appeared a year before the book's nominal publication date...)

The more I read about Nilo-Saharan, the less convinced I am that it exists (much less that Songhay belongs to it.) That means the classification of the languages of quite a lot of Africa is basically up for grabs. It would be great to have a reexamination of the area.