Sudanese Arabic dictionary and linguistic research resource, by James Dickins
Professor James Dickins
We are delighted to publish here on the website a longer version of Professor James Dickins’ article about Sudanese Arabic that recently appeared in our journal Sudan Studies. The article promotes online resources that are now freely available online.
1. History of the project
This article discusses an Arabic-English Sudanese Arabic dictionary and linguistic resource which I recently put up on the internet: https://www.phil.muni.cz/linguistica/dickins/. Both the dictionary and the linguistic resource are open for anyone to use (subject to the standard academic acknowledgement of the source of the material). I hope that they will be of use to learners of Sudanese Arabic as well as those with an academic research interest in the language.
In 1982, I began a Sudanese Arabic dictionary project with Robin Thelwall, at that time a Lecturer in Linguistics at the University of Ulster. I had previously taught English at Gezira Aba Higher Secondary School for Boys between 1980 and 1980, having studied Arabic at Cambridge University before that. The dictionary project began on card files but in 1984 I was able to transfer the material to a VAX mainframe computer at the University of Ulster. The main initial sources of the dictionary material were Hillelson (1930), Awn Alsharif Qasim’s (1977) Arabic-Arabic dictionary, and the vocabulary at the end of Persson and Persson (1980).
I subsequently returned to Sudan to do linguistic fieldwork in 1985 and 1986, and was able to check dictionary material with a friend of mine Ali Al-Rashid Al-Amin and other Sudanese who kindly helped me. The dictionary then lay dormant for many years, until I was able to revive it in the early 2000s when I was a lecturer at the University of Durham. During this period, I was able to transfer the dictionary material from a very large magnetic tape, via the University of Durham computer system to an Excel file, the format which it is still in today. I was then able to start adding further material to the dictionary.
In 2005, I was able to return to work on the dictionary more seriously, and in 2005 and 2006, working with a linguistic consultant, Elrayah Abdelgadir – then a lecturer in Arabic at the University of Marseilles – we were able to revise the entire dictionary. Since then, I have continued to work on the dictionary, and following my recent retirement, I hope to further develop it. I welcome collaboration with any other researchers who would like to join me in this project.
2. Outline of the dictionary and linguistic research resource
This is a downloadable Arabic-English – and, by extension English-Arabic – Sudanese Arabic dictionary and linguistic research resource. I will, for brevity, subsequently simply refer to it as Sudanese Arabic dictionary (or just ‘dictionary’, where this is appropriate). The dictionary is in Excel format and covers principally what I have elsewhere referred to as Central Urban Sudanese (e.g. Dickins 2023), i.e. the dialect of Khartoum and other urban areas of central Sudan. The information given in the Sudanese Arabic dictionary goes beyond traditional dictionary material, including, for example, also searchable and sortable morphological and semantic-field analyses. The dictionary can accordingly be used as a resource for linguistic research.
I thank the following, in particular, for their work on this project: Elrayah Abdelgadir (dictionary consultant, 2005 and 2007), Ali Al-Rashid Al-Amin (dictionary consultant, 1985, 1986), Ashraf Abdelhay, Asjad Saeed Awad Alhassan, Mohammed El Shazali, Taj Kandoura, Tarig Rahma, Yousif Elhindi, and Abd El Matallab Fahal. Many other people, not named here, have also contributed, and I thank them also. This is a work in progress, and , as noted, I intend to continue adding further to the dictionary.
If you have any comments on the dictionary or would like to contribute to it (by correcting errors, providing new words and phrases, etc.), or if you would like to discuss how to do linguistic research using the dictionary as a resource, please contact me on email: J.Dickins.leeds@outlook.com. You are also welcome to use the dictionary as a linguistic resource without contacting me. If you use this dictionary in your research, please reference it as:
Dickins, James. 2025. Sudanese Arabic dictionary and linguistic research resource. https://www.phil.muni.cz/linguistica/dickins/
3. Downloadable documents
All the downloadable documents are available from the following site: https://www.phil.muni.cz/linguistica/dickins/. They are as follows:
The Sudanese Arabic dictionary
The details of the transliteration system and other symbols (the ‘coding’) used for the dictionary.
The semantic-field classification adapted from the McArthur Lexicon of Contemporary English, used for the dictionary.
Extended Version of Rank Frequency List: Spoken English, from Frequencies in Spoken and Written English (Leech, Rayson and Wilson 2001) with semantic-field categories added for all entries.
All these documents are discussed in more detail in subsequent sections.
4. Sudanese Arabic dictionary
The dictionary attempts to cover the basic vocabulary of Sudanese Arabic by including Sudanese Arabic equivalents for almost all the vocabulary items found in Rank frequency list: spoken English from Word frequencies in written and spoken English (Leech, Rayson and Wilson 2001: 144-180). It also includes many other Sudanese Arabic words and phrases. As noted in Section 2, the dictionary can be downloaded from: https://www.phil.muni.cz/linguistica/dickins/
The dictionary is in Excel spreadsheet format, presented as an Arabic-English dictionary, ordered according to the traditional root system for Arabic dictionaries. For readers not familiar with the root system of Arabic, the following provides information: https://en.wikipedia.org/wiki/Semitic_root
Reflecting the fact that this is a work in progress, there is some repetition of material in the dictionary. At various points, too, question marks have been added to reflect the fact that some of the information (e.g. some English equivalents of Arabic entries) needs to be checked further.
The Excel format of the dictionary allows the reader to order the material in various ways (alphabetically on English entries, alphabetically on Arabic entries, by root, etc.). It also makes it possible for the reader to extract information of significance for linguistic research into Sudanese Arabic, e.g. morphological information or semantic-field information, and to cross-classify different kinds of information.
The dictionary contains the following information fields (columns):
Column A: Basic line number
Column B: English entry
Column C: Arabic entry 1 - basic
Column D: Arabic entry 2 - secondary
Column E: Word class (basic syntax, etc.)
Column F: Verb morphology
Column G: Non-verbal and deverbal morphology
Column H: Plural noun/adjective morphology
Column I: Style and regional usage
Column J: Semantic field
Column K: Root - 1: according to the traditional Arabic root system
Column L: Root - 2: roots with root letters converted to numbers, for sorting
These categories are discussed further in the next section.
5. Dictionary ‘coding’: transliteration system and other symbols
As noted in Section 2, the details of the transliteration system for Sudanese Arabic and the other ‘coding’ used in the dictionary can be downloaded from: https://www.phil.muni.cz/linguistica/dickins/ Note that spelling adopted for Sudanese Arabic forms is roughly a transliteration from Arabic rather than a transcription – i.e. the English forms correspond to Sudanese Arabic spellings (as normally found in dialect writing), rather than corresponding more directly to the phonology or phonetics of Sudanese Arabic.
The fact that I started working on this project on a VAX mainframe computer (https://en.wikipedia.org/wiki/VAX) in 1984, which had rather limited options in terms of line-length, column-length and available symbols explains some of the annotation decisions which I took.
6. Semantic-field categories
The semantic-field classification used for the dictionary (Column J) is adapted from the McArthur Lexicon of Contemporary English (McArthur 1991), and as noted in Section 2 can be downloaded from: https://www.phil.muni.cz/linguistica/dickins/. The Excel format of the semantic-field categorisation allows for sorting and extraction of information according to the following categories:
Column A: Primary category
Column B: Secondary category
Column C: Tertiary category
Column D: Coding of category
Column E: Word class
7. Extended Version of Rank Frequency List: Spoken English
This is a basic English vocabulary list taken from Frequencies in Spoken and Written English (Leech, Rayson and Wilson 2001), with the semantic-field categories in the McArthur Lexicon of contemporary English added for all list entries. It was used to try and cover the basic vocabulary of Sudanese Arabic in the dictionary. This list is based on the British National Corpus. Information about the corpus and electronic versions of frequency lists derived from it can be found at: http://ucrel.lancs.ac.uk/bncfreq/ As noted in Section 2, the Extended version of the rank frequency list: spoken English can be downloaded from: https://www.phil.muni.cz/linguistica/dickins/
The Excel format of the Extended version of the rank frequency list: spoken English allows for sorting and extraction of information according to the following categories:
Column A: Rank frequency order
Column B: Non-lemmatized head [as given in Leech, Rayson and Wilson 2001]
Column C: Lemmatized headword [i.e. standard dictionary-type headword]
Column D: McArthur category
Column E: Word-class
Column F: Rounded word frequency (per million word tokens) in speech
Column G: Log likelihood
Column H: Rounded word frequency (per million word tokens in writing)
Column I: Source of information (given throughout as WoFreSpoWriEng)
I thank Geoffrey Leech, Paul Rayson and Andrew Wilson, copyright holders of Word frequencies in spoken and written English, for permission to reproduce their material.
8. Using the Sudanese Arabic dictionary as an Arabic-English dictionary
The Sudanese Arabic dictionary can be used as an Arabic-English dictionary in two ways.
8.1 – by using the Arabic root to find Arabic headwords
If you know the root of the Arabic word, you can search on Column K ROOT -1: according to the traditional Arabic root system using the ‘text filter’ function, to look up this root. (For details of how to search on Excel using ‘text filter’, see: https://www.exceldemy.com/text-filter-in-excel/.) For the root <sm9, for example, this gives the following results:
8.2 – by using an individual Arabic word to find relevant Arabic entries
If you know the word simi9, for example, and want to find out its meanings, you can use Column C Arabic entry 1 to search for this word. This gives the following results:
Searching for a specific word in Column C may, as in this case, give some results in which the root listed in Column K, (ROOT -1: according to the traditional Arabic root system) is not the same as the root of the word which is searched for.
9 – Using the Sudanese Arabic dictionary as an English-Arabic dictionary
The simplest way to use the Sudanese Arabic dictionary as an English-Arabic dictionary is to search on Column B English entry. So, for example, searching on Column B for all entries which contain obey gives the following results:
Sometimes, however, this approach yields many irrelevant results. Take the case of looking for Sudanese Arabic equivalents of English hear. Searching on Column B for all entries which contain ‘hear’ (i.e. this string of letters) yields numerous results, the first twelve of which are:
None of these initial results are relevant; they all, in fact, involve heart, rather than hear (the string of letters ‘hear’ is, of course, found in heart as well as hear). There are various ways to get round this. One is to identify the major cause of the problem – in this case the word heart – and remove this from the search items by doing a search on Column B for all entries which contains hear but do not contain heart. This yields the following results:
This list is not without its problems. For instance, it includes the irrelevant words shear and shears. However, in the main the words in the list involve hear, and the relevant information can be easily found.
Another way of getting round having too many irrelevant results from searching for words which contain a string of letters, e.g. ‘hear’, in Column B English entry is to do this search first and then within it, do a second search in Column I Semantic field. So, if we first do a search on Column B English entry for all the entries which contain ‘hear’ and we then do a further search on Column I Semantic field for the McArthur Lexicon of contemporary English category [F272] hearing and listening, we get the following results, all of which are relevant:
10. Using the Sudanese Arabic dictionary as a thesaurus
To find words belonging to a particular semantic field, search in Column J Semantic field using the relevant McArthur Lexicon of contemporary English category. Thus, for example, the McArthur category [K055] amplifiers and microphones gives the following results:
11. Using the Sudanese Arabic dictionary as a research resource
There are many ways in which the Sudanese Arabic dictionary can be used as a research resource. I will consider one example in relation to the morphological pattern fa9laan. This is encoded in Column G Non-verbal and deverbal morphology as [401 ]. A search on Column G for this gives numerous results, of which the following are examples:
The standard plural for words on the fa9laan pattern in Sudanese Arabic is -iyn, notated as [f ] in Column H Plural noun/adjective morphology; talfaan ‘worthless’, plural talfaaniyn. There are some fa9laan words, however, which also have the plural fa9aala, notated as [061 ] in Column H. An example is ta9baan ‘tired, weary, exhausted’, plural ta9baaniyn or ta9aaba. Having searched for [401 ] in Column G Non-verbal and deverbal morphology, all the fa9laan form having the plural fa9aala can be found by then searching for [061 ] in Column H. This yields the following list:
In fact fa9aala plurals of fa9laan forms all seem to have a generalised and relatively abstract meaning. Thus ta9aaba people are not simply people who are tired, but those who have been worn down by life. sakaara (plural of sakraan ‘drunk’) are not simply drunk people but drunkards. kasaala (plural of kaslaan ‘lazy’) are people who are habitually lazy. (This information is not included in the current version of the dictionary, but I hope to include it in a later one.)
We can also look at fa9laan forms in terms of their semantic fields. The easiest way to do this is (having searched for all the fa9laan forms in Column G) to use Sort A-Z on Column J Semantic field. This shows that fa9laan forms have senses belonging to a number of semantic fields. For instance, the following fall under McArthur’s category [B096] Bodily states and associated activities - not showing energy:
The following fa9laan forms fall under McArthur’s category [F104] angry and annoyed:
The following fall under McArthur’s category [L173] ending things:
It would also be possible to investigate further aspects of fa9laan forms. For example, most of them have related fi9il verbs as can be seen by comparing them with verbs from the same roots (fi9il verbs are notated as [01i ] in Column G Verb morphology). It would also be possible to investigate the syntax of fa9laan forms by analysing the information for them in Column E Word class (basic syntax, etc.). I am happy to discuss possible research avenues with any researcher who would like to make use of this dictionary.
References
Dickins, J. 2023. ‘Definiteness, pronoun suffixes, genitives and two types of syntax in Sudanese Arabic’. In Journal of Semitic Studies. 68:2, pp. 679-712. https://doi.org/10.1093/jss/fgac035
Hillelson, S. 1930. Sudan Arabic: An English-Arabic vocabulary. London: Sudan Government.
Leech, G., Rayson, P., and Wilson, A. 2001. Frequencies in spoken and written English (Longman, 2001)
McArthur, T. 1991. Longman lexicon of contemporary English. London: Longman.
Persson, Andrew, and Persson, Janet P. with Hussein, Ahmad. [1979] 1980. Sudanese colloquial Arabic for beginnners. Horselys Green: Summer Institute of Linguistics.
Qasim, Awn Alsharif. 1977. qāmūs al-lahja l-ʕāmmiyya fi-s-sūdān (2n. Cairo: al-maktab al-miṣrī al-ḥadīṯ. (قاسم، عون الشريف. 1985. قاموس اللهجة العامية في السودان. القاهرة: المكتب المصري الحديث)
James Dickins was Professor of Arabic at the University of Leeds, UK until he retired in Sept. 2023. He has a BA in Arabic and Turkish from the Univer sity of Cambridge (1980) and a PhD in Arabic Linguistics from Heriot-Watt University (1990). He taught English at a government secondary school in Sudan from 1980 to 1982, and lived in Yemen between 2000 and 2001, and Oman between 2024 and 2025. He has taught Arabic and Arabic>English translation at the University of Cambridge, Heriot Watt University, and the universities of St. Andrews, Durham, Salford and Leeds. His publications in clude Extended Axiomatic Linguistics (1998), Standard Arabic: An Advanced Course (1999, with Janet Watson), Thinking Arabic Translation (2002; 2nd edition 2016, with Sandor Hervey and Ian Higgins), Sudanese Arabic: Phonematics and Syllable Structure (2007), and Thematic Structure and Para-syntax: Arabic as a Case Study (2020); and Sudanese Arabic Dictionary and Linguistic Resource: https://www.phil. muni.cz/linguistica/dickins/, Sudan Revolution Archive: Songs, Chants and Poems of the 2018-19 Sudanese Revolution (2024): https://sudanarchives.wordpress.com/, as well as. He is currently finishing off a book entitled Decline and Fall of the Case System and Construct State in Arabic.
This article, the Dictionary itself and related linguistic resources are also all available at Linguistica.