Sunday, March 31, 2013
Cree 1-key keyboards for Mac
Just a quick announcement that I've put up the Cree 1-key keyboard layouts for Mac.
Tuesday, February 19, 2013
Oji-Cree 1-Key Keyboard Layouts
ᐊᓂᔑᓂᓂᒧᐧᐃᐣ
The Oji-Cree 1-Key keyboard layouts are up for download. One of them (Oji-Cree 1) is the standard 1-Key Languagegeek layout. The other (K-Net) is a Unicode keyboard replication of an earlier Oji-Cree font. This hopefully will be useful for those who are already accustomed to this old font.
Both Windows and Mac keyboard layouts are available.
The Oji-Cree 1-Key keyboard layouts are up for download. One of them (Oji-Cree 1) is the standard 1-Key Languagegeek layout. The other (K-Net) is a Unicode keyboard replication of an earlier Oji-Cree font. This hopefully will be useful for those who are already accustomed to this old font.
Both Windows and Mac keyboard layouts are available.
Monday, February 18, 2013
Dene Syllabics Keyboards for Windows
After a busy few months away from Languagegeek, I found some time today to add a couple new keyboards: 1-Key Dene Syllabics for Windows. I’m going to keep working on 1-Key syllabics keyboards whenever I can.
Wednesday, November 21, 2012
Dene Syllabics Keyboards
It's been a long time since I updated the Dene Syllabics keyboard layouts. Well, I've gotten around to it today. There are two keyboard layouts up on the page: NWT Dene and Sayisi Dene.
NWT Dene
This follows the old Catholic Syllabics tradition and is notable in that it indicates:- Nasalization with a raised ˋ
- Aspirated K and the fricative X with a raised ᑦ
- Some ejective consonants with a raised ˈ
- Nasalization ˋ is distinguished from final K ᐠ
- Aspirated K and X ᑦ is distinguished from final M ᒼ
- Ejective ˈ is distinguished from final B ᑊ
- OskiDeneC - though I'll be making a better Oski font.
- Masinahikan
- Pitabek - Nice and traditional Dene design
Sayisi Dene
This follows the old Anglican Syllabics tradition and is notable in that it:- Differentiates the vowels U and O
- Has unique characters for the TL CH/SH clusters
- Can indicate more voicing distinctions
- OskiDeneA - though I'll be making a better Oski font.
- MasinahikanDene
Dakelh
Dakelh isn't ready yet, but I look forward to working on it soon.Tuesday, June 19, 2012
iPhone/iPad app: FV Chat
Happy days, our Chat app is up at the iTunes store!
Over the past year, I have been working with the FirstVoices Team to get out an app for speakers and learners of all the Native Languages in Canada (and beyond). The app works through Google and Facebook Chat systems - but we have added in keyboard support for over a hundred languages from coast to coast to coast, including syllabics. I've been using the App for Cree and Dene syllabics and have found it über useful. Oh, AND IT'S FREE!!!
Each Canadian Native language has its own layout, and where different dialects have different orthographies, multiple keyboards are available to choose from: Kwak̕wala and Kʷak̓ʷala are the same language, but written in two different systems, so there are two keyboard layouts. A generic Australian Aboriginal keyboard layout should cover most of the languages from Australia, there's a Māori layout, and I added in Diné Bizaad (Navajo) to help get people south of the boarder interested. Though once you have a look at the app, the interest generates itself!
iOS won't let me make system-wide keyboard layouts, so we have to do this from within an App, but if you want to use the keyboard in e-mail or something, you can copy/paste out of the Chat App, and there's a "Sandbox" area where you can type without actually chatting with someone.
As usual for Languagegeek, if your language isn't there, we'd like to add it. Please contact the FV team and we'll get started. There will be a one-time fee to get the keyboard programmed, but once it's in the app, users won't have to pay a cent.
Over the past year, I have been working with the FirstVoices Team to get out an app for speakers and learners of all the Native Languages in Canada (and beyond). The app works through Google and Facebook Chat systems - but we have added in keyboard support for over a hundred languages from coast to coast to coast, including syllabics. I've been using the App for Cree and Dene syllabics and have found it über useful. Oh, AND IT'S FREE!!!
Each Canadian Native language has its own layout, and where different dialects have different orthographies, multiple keyboards are available to choose from: Kwak̕wala and Kʷak̓ʷala are the same language, but written in two different systems, so there are two keyboard layouts. A generic Australian Aboriginal keyboard layout should cover most of the languages from Australia, there's a Māori layout, and I added in Diné Bizaad (Navajo) to help get people south of the boarder interested. Though once you have a look at the app, the interest generates itself!
iOS won't let me make system-wide keyboard layouts, so we have to do this from within an App, but if you want to use the keyboard in e-mail or something, you can copy/paste out of the Chat App, and there's a "Sandbox" area where you can type without actually chatting with someone.
As usual for Languagegeek, if your language isn't there, we'd like to add it. Please contact the FV team and we'll get started. There will be a one-time fee to get the keyboard programmed, but once it's in the app, users won't have to pay a cent.
Sunday, June 3, 2012
New Cree Keyboards
I’ve put up 4 new Cree keyboards. They’re of the “one key - one character” variety which I find myself liking more and more. The “type in Roman, get syllabic output” keyboards have ended up being clunkier than I’d like (though I’ll get around to fixing them) and they are only useful if you know Roman Orthography, which many speakers do not.
Similar keyboards for other languages are on the way.
Thinking back to last post... On second thought, I'll keep:
I tried out using the combining diacritics, and while it worked in some of my fonts, it decidedly did not work in the system font Euphemia. The dot diacritics ended up in bad places, the ring diacritics didn’t end up at all. Granted, Euphemia has not been updated to include the UCAS extended range.
So I'm back to using the on-top diacritics. It does make the keyboard design not quite as intuitive as I would have liked, the dot-on-top is a dead key so must be pressed first, before the base syllabic.
Similar keyboards for other languages are on the way.
Thinking back to last post... On second thought, I'll keep:
- the dot-on-top (for long vowels) ᐄ ᑑ ᑳ
- the ring-on-top (for Moose Cree y-finals). ᢱ ᢷ ᢸ
- and any other diacritics that are non-spacing
So I'm back to using the on-top diacritics. It does make the keyboard design not quite as intuitive as I would have liked, the dot-on-top is a dead key so must be pressed first, before the base syllabic.
Thursday, May 31, 2012
Canadian Aboriginal Syllabics Decombined
One look at the Unicode range Unified Canadian Aboriginal Syllabics (1400–14DF) shows that there are a lot of precomposed characters. For example, the syllabic for /ta/ ᑕ has several extensions: ᑖ, ᑡ, ᑢ, ᑣ, ᑤ
The guideline is that one should never use the character 1427 (CANADIAN SYLLABICS FINAL MIDDLE DOT) for the mid-dot diacritic, nor should one ever use a combining diacritic like 0307 (COMBINING DOT ABOVE) for the dot-on-top diacritic. Consequently, ᑖ 1456 (CANADIAN SYLLABICS TAA) and ᑕ̇ 1455 0307 (CANADIAN SYLLABICS TA + COMBINING DOT ABOVE) are not equivalent.
Ever since the addition of UCAS Extended, one can represent all the extended syllabic characters as single Unicode precomposed characters; at least one can for Cree, Ojibwe, Naskapi and Inuktitut. However, if base-syllabics with diacritics must be precomposed into a single character, then Blackfoot and Dene cannot be written following Unicode guidelines.
Before we carry on, there is an important principle to posit.
If we can’t use 1427 as a sort of diacritic, what is the point (pardon the pun) of this character anyway? The only source that I know of that would explain the inclusion of 1427 is the 1866 Recueil de prières, catéchisme, et cantiques à l’usage des Sauvages de la Baie d’Hudson (Moose Cree dialect). The writer follows the normal eastern practice of employing a left-side mid dot for /w/ onsets – on the chart on page 2, ᐺ is transcribed pwe. Interestingly for us, in the following row, ᐯᐧ is transcribed peu. Whereas most if not all other writers use a ring-final ᐤ (1424 CANADIAN SYLLABICS FINAL RING) to indicate a /w/ at the end of a syllable (in the coda), this version of the Recueil de prières uses a mid-dot.
This raises an encoding problem. As an example, let’s look on page 6 (dot-on-top marking is absent in the work), specifically the word ᓂ ᑕᐺᑕᐗᐧ /ni-tāpwētawāw/ I believe him. In other works, this would be written with a raised-ring final ᓂ ᑕᐺᑕᐗᐤ, but here final-w is written with a mid-dot.
The problem is, how can one know the difference between the mid-dots in ᓂ ᑕᐺᑕᐗᐧ ? The first and second mid-dots indicate the w-onset: ᐺ /pwē/, ᐗ /wā/ – but the last is a final-w. According to Unicode guidelines, the first two mid-dots are part of precomposed characters: ᐺ (143A) and ᐗ (1417). The final mid-dot would be ᐧ (1427). Here we have two different encodings of the same graphical symbol, which goes squarely against my policy as outlined in the previous post on apostrophes. One cannot expect typists to know that two otherwise identical symbols need two different encodings, it’s an unfair expectation.
So what to do for this variant of Moose Cree? The only thing to do is abandon the precomposed w-onset characters altogether, and use 1427 in all cases. Thus in the word ᓂ ᑕᐧᐯᑕᐧᐊᐧ
Barring any additions to Unicode to include these characters (which I’m not inclined to propose), the simple solution is to just use 1427 for all instances of s-onset in Blackfoot, no precomposed characters anywhere.
Yes there is, and it comes from Dene as written in the Northwest Territories. In fact there are two major issues.
However, in two very early Dënesųłiné versions of Prières, cantiques et catéchisme en langue montagnaise ou chipeweyan (1857, 1965), if the nasalized vowel had no consonant onset, that is it was a nasal version of the syllabics: ᐊ ᐁ ᐃ or ᐅ, the nasal diacritic went on top of the syllabic, ᐅ̀ᐣᐟᓯ /ǫntsi/. By 1890, this practice had ceased and the nasal diacritic always goes after the syllabic ᐅˋᐣᐟᓯ.
The UCAS Unicode range includes precomposed nasalized vowels, but only for the no-onset syllabics, ᐫ ᐬ ᐭ ᐮ – should the diacritic should go atop or to the right of the base? But any syllable can be nasalized in Dene, so we need to account for ᘕˋ /zǫ/, ᗃˋ /ghą/, ᓇˋ /ną/ and so on. To manage this, we have two choices, either make a proposal to add at least 64 new nasalized characters to Unicode, or use ˋ 02CB (MODIFIER LETTER GRAVE ACCENT) for all cases of the nasal diacritic. As a result, the precomposed characters ᐫ ᐬ ᐭ ᐮ should be abandoned completely, deprecrate the lot.
I suggest that this is both preferable and more economical. More economical because we don’t need over 64 new characters in Unicode, and preferable because the internal logic of the syllabics system seems to be that the nasal diacritic is distinct. In the Dëne Yatı Ɂerehtł'ı́s, one can see that nasalization can occur after a final: ᐁᐣᑯᕍᐩˋ /ekuląy/. Obviously we cannot compose the nasalization mark with the base character ᕍ /la/ with the final ᐩ /y/ in the way.
Thus we are forced to abandon a precomposed set of characters for a multi-characters series.
In one variant of Tłįchǫ syllabics, these unmarked ejectives are given the vertical tick: so ᐟᕃˈ is /tł'e/. To write this sequence, we need 02C8 (MODIFIER LETTER VERTICAL LINE).
The raised vertical tick is also used in Dënesųłiné to indicate aspiration: ˈᐁ is /he/: 02C8 1401.
We could propose to add 12 new characters to account for Tłįchǫ, plus another 4 for Dënesųłiné. Double this to include the nasalized variants (32) plus nasal versions of /t'/ and /k'/ is 40 new characters. But given the situation with nasal vowels, it seems easier to use 02C8 for all cases, so ᑕˈ is decomposed into 1455 02C8. So I’ll abandon the t' and k' series.
We gain little by using precomposed characters, after much reflection, I can’t see any benefit. My keyboard layouts have always used the precomposed characters, and I have received complaints from users:
I’m going to put out new Syllabics keyboards with 1427 mid-dots. This is what speakers have been asking for, and it will make everyone’s lives easier. It’s what the Cree Wikipedia does anyway!
- The dot-on-top diacritic signifies that the vowel is lengthened or more peripheral in some way: ᑖ is /tā/.
- A mid-dot diacritic to the left or right indicates that there is a /w/ glide in the onset of the syllable: ᑢ is /twa/. Whether the dot appears on the left or right of the base character depends on dialect/ortholect, so in most of Ontario and points east ᑡ is the accepted form for /twa/.
- Combine the two diacritics together and you get ᑤ or ᑣ ~ /twā/.
The guideline is that one should never use the character 1427 (CANADIAN SYLLABICS FINAL MIDDLE DOT) for the mid-dot diacritic, nor should one ever use a combining diacritic like 0307 (COMBINING DOT ABOVE) for the dot-on-top diacritic. Consequently, ᑖ 1456 (CANADIAN SYLLABICS TAA) and ᑕ̇ 1455 0307 (CANADIAN SYLLABICS TA + COMBINING DOT ABOVE) are not equivalent.
Ever since the addition of UCAS Extended, one can represent all the extended syllabic characters as single Unicode precomposed characters; at least one can for Cree, Ojibwe, Naskapi and Inuktitut. However, if base-syllabics with diacritics must be precomposed into a single character, then Blackfoot and Dene cannot be written following Unicode guidelines.
Before we carry on, there is an important principle to posit.
One symbol - one encoding Principle
As I discussed in the previous post, I believe that it is fundamentally a bad move to have one visual/graphical symbol with two different encodings and expect a typist to know the difference, and be dilligent about differentiating the two. For example, one could not expect an English typist to know that the apostrophe in won’t and the ‘single closing quotation mark’ are two different characters. Functionally they are quite different:- the apostrophe in won’t is lexical, it is built into the word and is essentially a letter.
- the ‘closing single quotation mark’ is punctuation, it is not in the dictionary, it is not part of the word mark.
Why have 1427 (CANADIAN SYLLABICS FINAL MIDDLE DOT) anyway?
Back to syllabics and the mid-dot diacritic.If we can’t use 1427 as a sort of diacritic, what is the point (pardon the pun) of this character anyway? The only source that I know of that would explain the inclusion of 1427 is the 1866 Recueil de prières, catéchisme, et cantiques à l’usage des Sauvages de la Baie d’Hudson (Moose Cree dialect). The writer follows the normal eastern practice of employing a left-side mid dot for /w/ onsets – on the chart on page 2, ᐺ is transcribed pwe. Interestingly for us, in the following row, ᐯᐧ is transcribed peu. Whereas most if not all other writers use a ring-final ᐤ (1424 CANADIAN SYLLABICS FINAL RING) to indicate a /w/ at the end of a syllable (in the coda), this version of the Recueil de prières uses a mid-dot.
This raises an encoding problem. As an example, let’s look on page 6 (dot-on-top marking is absent in the work), specifically the word ᓂ ᑕᐺᑕᐗᐧ /ni-tāpwētawāw/ I believe him. In other works, this would be written with a raised-ring final ᓂ ᑕᐺᑕᐗᐤ, but here final-w is written with a mid-dot.
The problem is, how can one know the difference between the mid-dots in ᓂ ᑕᐺᑕᐗᐧ ? The first and second mid-dots indicate the w-onset: ᐺ /pwē/, ᐗ /wā/ – but the last is a final-w. According to Unicode guidelines, the first two mid-dots are part of precomposed characters: ᐺ (143A) and ᐗ (1417). The final mid-dot would be ᐧ (1427). Here we have two different encodings of the same graphical symbol, which goes squarely against my policy as outlined in the previous post on apostrophes. One cannot expect typists to know that two otherwise identical symbols need two different encodings, it’s an unfair expectation.
So what to do for this variant of Moose Cree? The only thing to do is abandon the precomposed w-onset characters altogether, and use 1427 in all cases. Thus in the word ᓂ ᑕᐧᐯᑕᐧᐊᐧ
- the syllable /pwē/ ᐧᐯ should be the sequence 1427 142F
- the syllable /wāw/ᐧᐊᐧ should be 1427 140A 1427.
Blackfoot mid-dot
All right, so the above section was just about an obscure, historical orthographic variant of Moose Cree. Let’s move on to Blackfoot, where the mid-dot indicates an onset /s/: plain ᒣ is /ta/, ᒭ is /tsa/. Not a problem, one can use the precomposed ᒭ 14AD (CANADIAN SYLLABICS WEST-CREE MWE) to represent /tsa/. But Blackfoot can use the s-onset mid-dot after other syllabics base letters, most often the k-series: ᖿᐧ /ksa/. These Blackfoot k-series syllabics have no Unicode precomposed dotted characters.Barring any additions to Unicode to include these characters (which I’m not inclined to propose), the simple solution is to just use 1427 for all instances of s-onset in Blackfoot, no precomposed characters anywhere.
- ᖿᐧ is 15BF 1427
- ᒣᐧ is 14A3 1427 (not the precomposed 14AD)
Dene mid-dot
Dene has the same issue for w-onset:- ᐘ /wa/ : could use precomposed 1418
- ᑿ /gwa/ : could use precomposed 147F
- ᗃᐧ /ghwa/ : there is no precomposed character, must use 15C3 1427
- ᒈᐧ /k’wa/ : no precomposed character must use 1488 1427 (see later though for Dene ejectives)
Dene and why I have thrown in the towel
But is there something stopping the addition of these additional mid-dotted precomposed characters? Apart from the one symbol, two encodings problem for old Moose Cree?Yes there is, and it comes from Dene as written in the Northwest Territories. In fact there are two major issues.
Nasal vowels
In Dene, nasal vowels are indicated by a raised "Grave" diacritic. In virtually every written document in Dene, this diacritic is spacing, i.e. comes after the base glyph. So in a word like ᓀᘕˋ /nezǫ/ he/she is good, that final ˋ indicates that the vowel is nasal.However, in two very early Dënesųłiné versions of Prières, cantiques et catéchisme en langue montagnaise ou chipeweyan (1857, 1965), if the nasalized vowel had no consonant onset, that is it was a nasal version of the syllabics: ᐊ ᐁ ᐃ or ᐅ, the nasal diacritic went on top of the syllabic, ᐅ̀ᐣᐟᓯ /ǫntsi/. By 1890, this practice had ceased and the nasal diacritic always goes after the syllabic ᐅˋᐣᐟᓯ.
The UCAS Unicode range includes precomposed nasalized vowels, but only for the no-onset syllabics, ᐫ ᐬ ᐭ ᐮ – should the diacritic should go atop or to the right of the base? But any syllable can be nasalized in Dene, so we need to account for ᘕˋ /zǫ/, ᗃˋ /ghą/, ᓇˋ /ną/ and so on. To manage this, we have two choices, either make a proposal to add at least 64 new nasalized characters to Unicode, or use ˋ 02CB (MODIFIER LETTER GRAVE ACCENT) for all cases of the nasal diacritic. As a result, the precomposed characters ᐫ ᐬ ᐭ ᐮ should be abandoned completely, deprecrate the lot.
I suggest that this is both preferable and more economical. More economical because we don’t need over 64 new characters in Unicode, and preferable because the internal logic of the syllabics system seems to be that the nasal diacritic is distinct. In the Dëne Yatı Ɂerehtł'ı́s, one can see that nasalization can occur after a final: ᐁᐣᑯᕍᐩˋ /ekuląy/. Obviously we cannot compose the nasalization mark with the base character ᕍ /la/ with the final ᐩ /y/ in the way.
Thus we are forced to abandon a precomposed set of characters for a multi-characters series.
Ejective consonants
In Dene syllabics, ejective consonants are indicated with a raised vertical tick: ᑪ is /t'a/ (146A), ᒈ is /k'a/ (1488). In Dënesųłiné, there are other ejective consonants (ch', tł', ts') but these aren’t indicated in writing, and /tth'/ gets it’s own series ᕮ. No problem thus far.In one variant of Tłįchǫ syllabics, these unmarked ejectives are given the vertical tick: so ᐟᕃˈ is /tł'e/. To write this sequence, we need 02C8 (MODIFIER LETTER VERTICAL LINE).
The raised vertical tick is also used in Dënesųłiné to indicate aspiration: ˈᐁ is /he/: 02C8 1401.
We could propose to add 12 new characters to account for Tłįchǫ, plus another 4 for Dënesųłiné. Double this to include the nasalized variants (32) plus nasal versions of /t'/ and /k'/ is 40 new characters. But given the situation with nasal vowels, it seems easier to use 02C8 for all cases, so ᑕˈ is decomposed into 1455 02C8. So I’ll abandon the t' and k' series.
Conclusions
I could go on, and I’m sure I will in another post, but for now what have I concluded- For Old Moose Cree, precomposed mid-dot characters breaks the one symbol - one encoding principle. Thus all mid-dots in Old Moose Cree must be 1427 – all precomposed characters are disallowed.
- For Blackfoot and Dene to use only precomposed characters, we would need to add well over 100 new characters. If we abandon the precomposed characters, we can write these langauges with the UCAS as it is.
- We have to deprecate the nasalized-vowel series anyway because finals can interrupt the vowel-nasal bind.
We gain little by using precomposed characters, after much reflection, I can’t see any benefit. My keyboard layouts have always used the precomposed characters, and I have received complaints from users:
- They cannot just insert a mid-dot. In editing a document, the typist wants to correct a misspelled word, let’s say ᑕᐯ should be ᑕᐻ /tāpwē/. They want to simply put the cursor after the ᐯ, and type the dot-key. They can’t though and have to delete the ᐯ and retype ᐻ. This complaint indicates that conceptually the mid-dot is a discrete entity.
- Dot-on-top vowel indication is optional in many dialects. For /tāpwē/, it is acceptable for many to write either ᑖᐻ or ᑕᐻ. With precomposed characters, ᑕ (1455) and ᑖ (1456) are completely different characters, whereas if we have a common base character ᑕ which may or may not be followed by a combining accent, it is simple to ignore the accent mark and treat ᑖᐻ and ᑕᐻ as the same word in searches or spell checkers. People see the dotted and undotted characters as the same underlying thing, with one perhaps a more precise spelling.
What am I to do?
I’ve thrown in the towel and am not going to bother with precomposed syllabics any longer. They’re gone, out of my life.I’m going to put out new Syllabics keyboards with 1427 mid-dots. This is what speakers have been asking for, and it will make everyone’s lives easier. It’s what the Cree Wikipedia does anyway!
Subscribe to:
Posts (Atom)
