Monday, February 9, 2009

how phonetic is English? Thoughts.

http://www.spellingsociety.org/journals/j32/dewey.php
The figure of 85% regular is often quoted as if it were based on solid research [Crystal, 1999]. The original research was done by Dewey in the 1940's and repeated by Paul Hanna in the 1960's. Hanna noted that you can guess with 75% accuracy the dictionary spelling for each phoneme with 4 guesses (see www.unifon.org/uu-29ways.html]. Predicting phoneme spelling is not the same as predicting syllable spelling [see chart]. When people, such as Flesch [1956, 1983], say that English has a highly regular orthography or is 97% phonemic, they have something else in mind other than predictability. Spaulding (1964] uses 70+ phonograms and 26 exception rules to arrive at her high estimate for English regularity. With around 200 sequentially applied exception rules and two spellings per sound, traditionally spelled words can be shown to have a high degree of predictability. Memorizing 200 rules, however, might prove to be more difficult for humans than memorizing the dictionary

D: I am reminded of the claims about the # of rules needed for grammar.
English: 1000(s).
Optimal IAL: about 2 dozen. I imagine a score is doable.

We notice a parallel in the # of rules needed to spell correctly.
English spelling is more erratic than many other natural language spelling conventions.
Spanish is known to be highly regular. This is in part due to historical conscious reforms. Important, since this shows that a con-lang (of sorts) can be successful and achieve a desirable and desired goal. If it can work for a con-lang - maybe an aux-lang too?

How Phonemic depends on the unit of analysis

PhonemesSyllablesWords
75% regular50% regular40% regular


Thus, the current edition of How we spell!, [1] formally English Heterography, identifies in a single [72,000 word], abridged dictionary 530 spellings of 41 sounds, employing 273 different symbols, that is 12.9 graphemes per phoneme, 1.9 phonemes per grapheme.

D: about 13 ways to show each sound, and about 2 sounds per letter.
Many folks speak of learning failures, when students with particular hardships encounter difficulty. I speak of teaching failures, in the sense that the subject matter is unnecessarily far too complex to meet its goal. Call me naive, but the goal of a language should be successful communication.

Aside: I had an insight yesterday. Fraternities and sports teams have historically used hazing rituals. Social psychology suggests members would then value their membership more due to the effort required to obtain them. Nobody wants to admit that one did (following humiliating behavior) for an insignificant goal!
Similarly, I have noticed a trend with the literati of English, as well as all the classics majors I know. The amount of effort required for an English speaker to acquire functional Latin in adulthood is prodigious. Their mental endowments are indeed impressive. My friend "M" is adamantly against designed or reformed languages. She is a fluent Latin speaker. I wonder how much she may be annoyed that one could learn all the benefits of a (not yet born new) language in a tiny fraction of the time required to obtain a presently dead one. In the case of English-lovers, the present living de facto standard one.

"Thus, Laubach, [8] whose extraordinary achievements, "Each one teach one," in promoting literacy in over 300 languages thruout the world are well-known, employs for English a notation of 96 symbols [9] - actually, counting 4 recent additions and 18 doubled consonants, 118 symbols - several of them involving a diacritic, the macron; and describes as "regular" all spellings within the compass of that notation. "

D: there are sometimes multiple physical ways to make the same sound. Alternatively, sounds that we consider indistinguishable may be be considered meaningful in other languages.
I have pondered this with Hioxian.
Update: I don't need to show HIOX's tongue position nearly as much as I once thought.
Along the top of the HIOX symbols, the various vertical and diagonal bar segments stand for
1) lip
2) teeth
3) velar ridge
4) hard palate
5) soft palate.
I realized the appropriate tongue part adjacent to the above parts automatically is used. I do not need to show tongue tip/mid/rear positions. It is redundant.
However, such features as lip rounding and tongue narrowing may still matter.
I suspect the diacritic will be used for various instructions for forming a sound not obvious in a 2D cross-section, only apparent from the front. As well, I suspect amount of air flow and type, as well as duration (gemination) will be indicated on the diacritic.
I am not sure yet, but I hope the diacritic will remain optional in typical daily use.
So far, the consonant top/right/vertical bar is reserved for nasal, the bottom/right for voiced.
Voiced is important enough that I do not wish to assign it to a diacritic.
The visual appearnance and theme of characters is important.
For example, my Decimese proposed CVCVCV...(nasal consonant) will have the following format
1) first character, overt indication of voiced versus voiceless (or v.v.)
2) a middle-word consonant, the opposite
3) vocabularly item word termination, nasal consonant, top/right/vertical bar segment.
Combined with the 2 horizontal in-line diacritic bar segments indicating duration, there is a clear and obvious set of instruction for forming the word.
I know it is not strictly true, but indicating syllable stress via duration is workable. It is not true of many langauges, however, such as French.
English: FIVE, TEN, FIFteen, TWENty.
French words would have the same duration per syllable, as would New Zealand dialect English.

Another indication HIOXian should portray is the difference between a adjacent vowels and diphthongs. The HIOX vowel system if fairly straightforward. See the vowel chart in IPA. Hybrid 'mid' position vowels can be indicated via both adjacent location bar segments.
For that matter, I still have access to additional unused bar segments.
I would very much like to avoid using the vowel diacritic vertical and diagonal bar segments initially. This gives me the option to portray tones.

Aside: I took literacy tutor training this weekend for the local reading group.
We learned about various brain differences that affect written language acquisition.
An example of visual processing problems was portrayed as a chaotic mess of letters in a written note. The lines were not in line, and the class was stumped by what it meant.
We were further stymied by various letters being inverted or otherwise rotated.
Technically, HIOX will be vulnerable to this. I must assume that at least a few characters will still be intelligible when flipped horizonatally, vertically or diagonally.
I believe that the ideographic nature of the HIOX character will prevent this. However, dyslexia has left-right directionality issues. At first, the standard left-facing articulation diagram might not be clearly remembered.
Ygyde did a good job of addressing this.

----
English quirk of the day:

Ways to make various vowel sounds.

http://www.say-it-in-english.com/SpellHome.html
Why does the English language have so many words that are difficult to spell? The main reason is that English has 1,100 different ways to spell its 44 separate sounds, more than any other language.

Ways to spell Long 'U' shoe, grew, through. do, doom, flue, two, who, brute, duty
Ways to spell Long 'O' go, show, though, sew, beau, float, bone,
Ways to spell Long 'A' may, weigh, late, pain, rein, great
Ways to spell Long 'E' free, bean, magazine, gene, mete, be, mien, receive, believe
Ways to spell Long 'I' fine, rhyme, fight, align, isometric, bayou

Thursday, February 5, 2009

Punctuation. Are there rules? Or is it a matter of personal style?



http://www.theglobeandmail.com/servlet/story/RTGAM.20090205.wkoringdiscussion0205/BNStory/International/home

" 'We are not going to be able to rebuild Afghanistan into a Jeffersonian democracy,' " he said recently in a sharp and sweeping contrast to his predecessor."

Wow. I took me a while to parse that in my head.
Nestled punctuation, indeed!

From Wikipedia:

Usage

Quotations and speech

Single or double quotation marks denote either speech or a quotation. Neither style – single or double – is an absolute rule, though double quotation marks are preferred in the United States, and both single and double quotation marks are used in the United Kingdom. A publisher’s or even an author’s style may take precedence over national general preferences.

Oh, the irony! The same day in the same paper has the following article too.

http://www.theglobeandmail.com/servlet/story/RTGAM.20090204.wrussell0205/BNStory/Entertainment/home

Punctuation as ideological warfare



According to Truss's book, punctuation seems to have between 3 and 8 functions per symbol.
In the strictest sense, more like 3 to 5.
I would like to supplement the HIOXian letter system with revised punctuation.
It simply provides the option to overtly indicate via modified diacritic which function a punctuation symbol is fulfilling.
One can just use a comma. Or one can indicate the comma is indicating a list.

Hioxian punctuation will also be clearly left-and-right pairs for various symbols, such as parantheses and quotation marks. Keep in mind that no HIOXian figure uses any element smaller than a bar segment. There are no 'dots'. Dots are easy to miss, being tiny. I suspect
half the reason folks are using more dashes these days is for visual clarity.

. ; : , - - - -

As the Baby Boomers age, visual clarity for reading a computer display will matter more.

I personally use a nifty Firefox plug-in called Image Zoom.

One can increase text size via Control + mouse scroll wheel.
Oddly, my WinXP changes font size but Macs simply zoom as if the text is an image.
The result on a Mac is a large but fuzzy letter. Not much better...


I really do need to attach some tabs to key passages in her book.
Her humour and wit was much appreciated.
Gawd, she'd HATE my blog, LOL.

Wednesday, February 4, 2009

apologies for truncated blog yesterday. font designers, 3 iterations for decimese?

D: this blog builder does not fare well with cut and paste.
If the pasted section is wider than the permitted screen,
things get cut off.
Many small font passages become invisible.
When I catch them, I change their colour.
Otherwise, just highlight with "select all".

Good news! My pal Sanjay ordered me a font designer.
We shopped around and found that Fontlab's stuff is both
as good as any and cheaper. 100 bux bought me even the
ability to make Adobe-compatible fonts.
I did find the the SIB bitmap font maker for 30 bux, but it is
pretty limited. Many other programs could handle OTF and
TTF but not Adobe.

I can now roll-my-own Old Persian font.
I can also make a working font for my HIOXian system.

Next stop: I'm building a 3d polarized rig this month. It involves
2 cruddy old used LCD monitors from the local campus surplus
sale, a half-silvered mirror ($40-100) and a nice wood box.
I am intensely curious about loading language information on top
of existing standard alphabet writing.
A local university computer guy and another computer prof too
are both pitching in to get things going.
I am also am pals with one U computer hardware gal, so I have
my bases covered.
( Humming tune of "I get by with a little help from my friends...".)
[=

Hmm. I wonder if I could show pitch with depth information...

Well, if nothing else, my display will make a sweet gaming rig for
World of Warcraft, LOL.

Re: HIOXian font.
1) standard version, medium thick bar segments, straight and diagonal lines only
2) baroque style, circular/curved/swirly motif
3) examples of cursive handwriting, also a font.
D: I don't think any existing software can accommodate my colour-coded
Rainbowesque letter-stacking idea.
There is no reason that could not also work with existing fonts.
It would work better, I think, with letters that overlap heavily.

I'll likely build HIOXian to the 2 suggested standards of LangX.

http://bahai-library.com/books/lango/lang02.html#%

I will need to think about keyboard layout. I looked into that before.
I have an idea for a "hand and a half" keyboard.
It emphasizes only dominant hand dexterity.

A coupla thoughts on HIOXian.
1) keyboard layout needs to be based on QWERTY and DVORAK
considerations of letter frequency
2) the order of the 'alphabet' will likely be an homage to the Indian
system which methodically lists letters by place of articulation.

http://www.omniglot.com/writing/brahmi.htm

http://en.wikipedia.org/wiki/Alphabet#Middle_Eastern_Scripts

Another homage, readily apparent in HIOXian:
In Korea, the Hangeul alphabet was scientifically created by Korean scholars under King Sejong in 1443. Understanding of phonetic alphabet of Mongolian Phagspa script aided in creation of phonetic script that suited Korean vocal language. Mongolian Phagspa script in turn was derived from the Brahmi script. Hangeul is a unique alphabet in a variety of ways: it is a featural alphabet, where many of the letters are designed off of a sound's place of articulation (P to look like widened mouth, L sound to look like tongue pulled in, etc.); it was consciously designed by the government at the time; and it situates individual letters into syllable clusters with equal dimensions as Chinese characters to allow for mixed script writing (one syllable always takes up one type-space no matter how many letters get stacked into building that one sound-block).

D: from LangX.

A Consonantal Script

The potential print-saving achievable by a consonantal script is astounding. With 27 consonants, 551,880 words of four letters or less are possible (27 + [27 × 27 =] 729 + [27 × 729 =] 19,683 + [27 × 19,683 =] 531,441 = 551,880) - four or fives times more than the total vocabulary of English (if the endless progression of names for numbers, chemical compounds etc. is excluded).

D: using the single digit numbers to denote which word could work.

This would work better earlier, with fewer vowels.

E.g. CVCV(nasal ending). V1 and V2 , if there are 5 each, allows 25 possibilities.

... OK I just refuted myself. There is no way to shorten the word this way.

The only benefit would one could assign vowels to the shift-for-capital letter category.

The #s could be briefer.

Language X is onto something important. One must adhere to standard computer hardware.
Read about the hassle of trying to type in Chinese!
Better yet, find out how hard a Chinese dictionary is to use.
The keyboard pretty much dictates a cap of c. 52 phonemes/letters in the near future.
Again, beyond 100 years from now we have NO idea what things will look like.
Perhaps the QWERTY keyboard will be supplanted by direct brain implants. Who knows?
A game plan past 500 years from now assuming business-as-usual is very conservative.

If Kurzweil is right, even planning beyond one generation has serious issues.

D: I attempted to try a LangX approach to Decimese.
Generation 1: 5 consonants pairs, 3 vowels, generic nasal ending.
Generation 2: 5, 5, 2 nasals
Generation 3: 6, 5 plus diphthongs, 3 nasals.
D: a carefully planned offloading of word particles onto consonant clusters, new sounds, as
well as diphthongs could result in very rapidly spoken synthetic language.
Generation 1 is likely a decent basis for a starter IAL, with world appeal.
Generation 2 is within the reach of Mandarin speakers.
Generation 3 requires Cantonese or more.
If one is in love with the idea of a highly variable word order, then Gen3 would do so.
I.e. M, N, NG word termination for noun/verb vocabulary items.
Once subject, object (nouns) and adjective/adverbs and verbs are addressed, one can engage
in Latin-like variable word order.
My friend in classics explained Latin to me once. I was gobsmacked by the complexity.

Generation examples.
1 - PB (same thing for now) - AUI (pick one) - MNNG (generic nasal ending)
2 - as per 1 but SH/CH pair, WY or LR consonant cluster, AUIEO, M or N
3 - as per but W Y L R, (vowel diphthongs), M N NG. H is now a consonant.
Note how limited Generation 3 is compared to Lang53 of LangX.
This makes it a near-future proposition.

Words in each generation.
1) pb aui m, e.g. pam
2) p lr a m
3) p l a y m.

D: a quick scan of the dictionary will indicate what consonant clusters are acceptable
to English speakers.
A list of Chinese monosyllables will the same for them.

In this respect, Decimese is unlike most IALs other than Ceqli.
It serves Mandarin interests now, and English interests later.
Generation 1 is the nod to the needs of the rest of the world.
Vocabulary items are culturally neutral though.
In that respect, it resembles LangX.
Ceqli has too many phonemes from the very start to serve as a world IAL.
At most, it could serve as a Chinese-English interlang.
Which is what it is designed for, as well as a valid design purpose.

Well enough rambling thoughts.
I'll try to stick to 1 topic tomorrow.

Tuesday, February 3, 2009

computer speech translation, a computer interlingua

http://www.newscientist.com/article/dn16528-innovation-speech-prediction-software.html

Word perfect

"Around the turn of this century it really was appalling – IBM's ViaVoice at that time seemed to have a 90% error rate. But nine years later Dragon Naturally Speaking 10 is for me more than 95% accurate. And when it's wrong, it's usually my fault.

AIST's idea could make such software even more powerful, by increasing the speed and accuracy with which you can dictate long and difficult words and common phrases."

D: the idea that we need to make auditorially distinct sounds for the computer is seeming antiquated.

However, the idea of a grammatical structure that parses easily remains valid.

I read the VOS is the clearest for computers. Note this is also what human pantomime suggests is the most intuitive for humans. (If you want references to websites and studies, I suggest you peruse my older blogs.)

I found a very thoughtful computer interlingua proposal. It is called "lexical semantics of a machine translation interlingua".

http://www.eskimo.com/~ram/lexical_semantics.html

Note how closely it resembles a modern Creole-esque IAL.

D: section 25 has useful design principles.

With the above in mind, we can state several general guidelines for word design:

    "1. Start with simple, common verbs and adjectives.  Isolate their
root concepts and apply it to every classifier. Appropriate suffixes
should be used when related verbs have different argument structures
(e.g. "to say" vs. "to tell"). In the process, a very large number
of less common concepts will be automatically derived. This
principle also applies to numeric, deictic, tense-aspect, and modal
concepts.

2. Keep in mind the inherent difference between basic state concepts
and modal concepts. When in doubt, always test new concepts to
determine if they are modal.

3. If there's difficulty defining a basic state or modality, or if
it has limited usefulness when combined with most classifiers, it is
very likely that the state is not very basic. When this occurs,
postpone derivation until later. You may be able to "accidentally"
derive it from a different root.

4. Always be suspicious of roots that represent energetic states.
Many of these concepts can actually be derived from non-energetic
states that end up being much more productive."

D: see their proposal for kin relationships.
It is well thought out (Section 25.4).

You will note a distinct taxonomic trend in vocabulary design, but without
the excess associated with languages by Wilkin and that of Ro.

D: this site tremendous potential for designing the core vocabulary for a
human interlang.
Being arbitrarily computer optimized, it is necessarily cultural neutral.
It is much more methodical, however.
The main problem with a taxonomic language design has always been that only
one minimal pair is present to prevent misunderstanding. Context is NOT an
aid, since the two words sound so similar in the same category.
Essentially, the trade off becomes
1) easier to learn, can guess general meaning but
2) less clear once in use at colloquial speeds.
Again, my HIOXian letter system should point out phoneme combinations that will
cause particular problems.

D: the emphasis on reducing the number of primitives (basic morphemes) needed
inspired me. I applied that tactic with closed class "function words" in English.
It will form the basis for the "function words" of Decimese (why am I calling it
that still when I have more than 5 consonant pairs now?).
For example, English has the pronouns I, we, you (plural implied), he/she/it and they.
Esperanto touches upon the idea of modular pronouns with the pair of il and ili.
If we parse about English pronouns, we end up with the following core concepts:
1) distance. inside, close and far ( a distinction of some languages )
2) quantity. single and plural.
3) gender, with neuter, masculine and feminine.
Well, why not build the pronouns in modular fashion from these concepts?
Decimese attempts to mitigate the one-minimal-pair clarity issue of taxonomic design
by using the syllable of CV, not C or V letter as the core unit.
This, in turn, requires shorthand version to address the issue of lost brevity.
Get something, lose something. The challenge is to finesse the concepts so the overall
language is more than a zero-sum-game of design elements, where all seem to be a
comparable set of pros and cons.
In the case of pronouns, a taxonomic system, or even a compound -concept approach
would require the brevity of the one grapheme to one phoneme taxonomy.

I think I'd be willing to confess that Esperanto, in its spotty and hazy and sporadic
fashion has touched upon most if not all clever language innovations.

If I stand taller one day, it will only be since I stand on the shoulders of that
giant Zamenhof.

(Out of time, not proofread. Apologies.)




Monday, February 2, 2009

IAL phoneme selection

D: presumably, an aux-lang (auxiliary language) is
1) global
2) acquired in part by adults who
3) are not academics with that much time or inclination.
This means the sounds need to be pretty common.
Back to my observation that just matching sounds to regional interlanguages may be adequate.

From Rick Harrison:

http://www.rickharrison.com/language/farewell.html
D: this is hard to read but likely right. To date, most folks consider an IAL a solution to a problem that does not exist.

http://www.rickharrison.com/language/optimal.html
D: this may very well be THE single document to read on IALs.

"1. An optimal IAL will be relatively easy for most children and adults to learn as a second language.

2. An optimal IAL will have the ability to handle both mundane conversation and highly technical information.

3. An optimal IAL will be culturally neutral; it will not provide advantages of word recognition or other special favors to one or two ethnic groups at the expense of all others."

D: I disagree generally with the culturally neutral requirement. Right now, getting English-speakers on board is essential. Yup, it is shameless sucking up the powers that be. In a generation, that power (economically, somewhat) will be Mandarin Chinese.

The requirement to be culturally neutral is one more lofty principle at the expense of a practical sales job.

"Morneau surveyed data on 25 major languages and indicated that the following phonemes are used in at least 22 of the 25: /a, e, i, o, u, b, d, k, l, m, n, p, s, t, y/ (“y” represents the semi-vowel...
Sapir et al. recommended an even smaller array of phonemes: /a, i, u, p, t, k, s, l, m, n, v/.
The most unmarked phonemes would be these: /a, i, u, p, t, k, m, n, s, l/. A second rank of slightly more marked, but still generally manageable phonemes would be: /e, o, b, d, g, f, h, y, w/. A third rank of dubious but possible phonemes would be: /v, z, r, ch, sh/.”

D: I will mentions some languages that are not meant to be modern IALs. I realize this is unfair.
Toki Pona just wanted to explore Taoist philosophy. Ceqli is meant to appease the Chinese. Esperanto is a century old, before complete data sets were available.
English and Mandarin are control groups.

http://en.wikipedia.org/wiki/Toki_Pona

Toki Pona has nine consonants (/p, t, k, s, m, n, l, j, w/) and five vowels (/a, e, i, o, u/). The first syllable of a word is stressed;[9] an initial vowel may be optionally proceeded by a glottal stop.[10] There are no diphthongs or long vowels, no consonant clusters, and no tone.

D: a coupla comments.
First, the consonant choices are nearly ideal. Voice/voiceless distinctions are not used. Notably, TP would be better off with only 3 vowels for near-universality. (Arabic has aui and AUI).

Ygyde:http://www.medianet.pl/~andrew/ygyde/ygyde.htm

6 vowels a e y o u i, 15 consonants b p d t g k w f z s j c m n l...
D: this is a fairly mechanistic approach. Note that most voiced/voiceless pairs are present, though V is absent based on rarity. So too is R missing due to rarity.
(I am not sure what Sapir based his selection upon.)
This is a sensible tier 1 and 2 phoneme selection. His use of H regarding voiced/voiceless consonant pairs is highly viable.

"The phoneme h is placed before the three letter long word only if its second phoneme is unvoiced consonant or n: p, t, k, f, s, c, n"

Ceqli is best understood as an interlang between English and Mandarin.
As such, it has too many sounds for a world IAL.

http://www.geocities.com/ceqli/alph.html

The Ceqli language uses the 26 letters of the Roman alphabet. 19 consonants:
And five vowels:
And two semivowels:

W and Y make these diphthongs: (11).

Mandarin:

http://en.wikipedia.org/wiki/Standard_Mandarin#Phonology
D: although it uses many phonemes, it has restrictive rules about syllable formation.
The number of syllables is considerably less than would otherwise be the case.

We all know English. It also has many phonemes and permissive syllable construction rules.
The word "strengths" is a good example.
I suspect the closest a Mandarin speaker could handle is "taron" or thereabouts.

And finally, we come to powder keg. Wow, those Esperantists certainly are vocal.
See the start of this entry. It is too complex as an adult, second-language, non-academic and global IAL. Not too bad as a European interlang.
Reviews of Esp-o indicate it is difficult for Asian language speakers. At the intermediate level, the level of infixing and agreement is difficult for English speakers.
I suspect a combination of too many phonemes and lax syllable construction rules are to blame.
(Say s-t-s much? Scii...)

http://en.wikipedia.org/wiki/Esperanto_phonology

A syllable in Esperanto is generally of the form (s/ŝ)(C)(C)V(C)(C).
Geminate consonants generally only occur in polymorphemic words, such as mal-longa "short", ek-kuŝi "to flop down", mis-skribi "to mis-write";...
(D: plus i-i adjacent. Again, ill thought out side-by-side syllable combinations.)

The phonemic inventory is essentially Slavic, as is much of the semantics, ... Esperanto has 22 consonants, 5 vowels, and two semivowels... (and diphthongs too).

D: the idea of a grapheme where one letter is assigned one sound only is nice. But using all 26 Roman alphabet letters immediately puts the phoneme inventory beyond the collective world.

Language X attempts to address this. Limited to a small phoneme selection such as Sapir suggested (a, i, u, p, t, k, s, l, m, n, v) results in very few syllables permitted. They wish to gradually add more phonemes once LangX is the mother tongue to the world.
Eventually they wish to have c. 26 consonants and vowels. The QWERTY keyboard pretty much dictates this limit. However, the assumption that we will still use QWERTY keyboards in 500 years seems conservative.

D: A 6-7 vowel AND consonant combination in 4 stages has some benefits. With only 13 sounds at stage 1, we do get stuck missing some very common consonants while using some less common vowels. But at stage 2 we have 13C/13V and by stage 4 we have 26C/26V.
I had pondered the option of cross-mapping these stages into tone pitches.
I.e. stages 1 to 4 would be 4 to 16 tones.
I.e. do mi so ti (4 whole notes), do re mi fa so la ti do (1+ octave), 12 (half notes), 16 (either beyond one octave, which requires training, or quarter notes, which still requires exposure during upbringing).

The objection that most world languages do not use tone for lexical purposes only applies prior to LangX as the universal world mother tongue.
For the same reason LangX is willing to become more synthetic with affixes, as well as with many more phonemes and presumably loosened syllable construction rules later on, tone is also legitimate to use.
I see an interesting convergence between LangX's Lang53 and Heinlen's SpeedTalk fictional language. The brevity of speech with 53 phonemes and broad syllable construction rules would be great. An interesting approach might be to use consonant clusters to reincorporate word particles into a more affixing, synthetic approach. The benefit of this is that a consonant cluster, unlike an extra syllable, does not increase speaking time much.
This is what I'd like to do with Decimese.

What have I learned from this entry?
I simply cannot use 10 vowels. The l/r distinction cannot presently be used.

http://www.acoustics.org/press/143rd/Guenther.html
D: it shows a Japanese speaker's brain processing L and R as the same sound.

D: I also believe we cannot use the voiced/voiceless consonant distinction arbitrarily.

Catering to Mandarin speakers means fewer syllable options. Looking at how syllables side-by-side will be pronounced is important.
Mandarin syllable forms are further constrained by English phoneme limitations.
There won't be very many one syllable words.
Consonant clusters must be approached cautiously and preferably optionally.
Use of word particles or the option of an affix in the form of a consonant cluster would maintain brevity.
My -N -M -NG syllable termination selection is not viable. ( I think it would be in Cantonese.)
I confess Decimeseis more an English-Chinese compromise with a few nods to the world.
OK, so how do I make use of just the -N and -NG endings for most vocabulary items?
Modify word order.
English: SVO. Article-adjective-noun, verb-adverb, (as per subject for object).
I tend to write that sSVvoO.
The problem is how the endings will look with -N and -NG.
ng N N ng ng O.
D: having adjacent identical endings prevent word order form indicating grammatical function clearly. I wish to have word roots with ambiguous grammatical function. The grammar role is to be indicated via a combination of word order AND word ending.
This can be accomplished with word order sSvVoO.
I am afraid we need to "boldly go", not "go boldly".
But if it is good enough for Kirk then who am I to argue?

[=
http://www.rishon-rishon.com/archives/096770.php

Mandarin Syllables

Cool chart of Mandarin syllables. There are 22 initials and 35 finals. Add the four tones, and you get a theoretical maximum of 3080 possible syllables. But, as you can see from the chart, only about two-thirds of the possibilities are actually used, which means that Mandarin uses only around 2000 syllables (I didn't count).

http://wiki.answers.com/Q/How_many_words_have_more_than_one_syllable
Dr P E Fischer Ph.D. lists over 3,500 one-syllable English words in her word-list book written for teachers of English to children.

Using 3,500 as a conservative estimate of one syllable words, there means there would be over 612,000 words of two or more syllables. But even if we double that figure arbitrarily to 7,000 one-syllable words,

Friday, January 30, 2009

as easy as 1-2-3. #s in English are a drag on economy

http://chineseculture.about.com/b/2008/12/16/why-are-chinese-better-at-math.htm

Author Malcolm Gladwell has a history of thought-provoking insight in his previous books,The Tipping Point and Blink. Now he's tackling another fascinating subject: success, in his latest work, Outliers.

D: my roomie mentioned this.

(Chapter one of author's book.)

D: Regarding memorizing a list of numbers.

Take a look at the following list of numbers: 4, 8, 5, 3, 9, 7, 6. Read them out loud. Now look away and spend twenty seconds memorizing that sequence before saying them out loud again. If you speak English, you have about a 50 percent chance of remembering that sequence perfectly. If you're Chinese, though, you're almost certain to get it right every time." The reason behind this, Gladwell writes, is because humans can store digits in a memory loop that last only about two seconds. In Chinese languages, numbers are shorter, allowing Chinese to both speak and remember those numbers in two seconds -- a fraction of the time it takes to remember those numbers in English.

D: regarding counting systems.

Moreover, Asian languages such as Chinese, Japanese and Korean have a more logical counting system compared to the irregular ways that numerals are spoken in English. As Gladwell writes: Eleven is ten-one (十一 in Chinese), twelve is ten-two (十二) and thirteen is ten-three (十三) and so on.

D: regarding fractions.

Even fractions are easier for Asian children because they are more easily understood and conceptual. For example one-half (fifty percent) is understood as 百分之五十 (bǎi fēn zhī shí) or literally, fifty parts out of 100 parts...

D: this is a common area of improvement for an IAL.

The #s 2-9 are expressed via C1to9+A. I.e. 2 ba, 3 da, 4, cha(or ca), 5 la, 6 ra, 7 tha(or ta), 8 va, and 9 wa.

D: that is from my Deafese first language attempt.

-------------
D: I love the saying "as easy as 1 2 3".
Cuz it's NOT.
Witness this.
1 2 3
Spelling: one, two, three
Consonant/vowel: VCV, CCV, CCCVV
Phonetically: wun, tu, TrE
C/V: CVC, CV,CCV

D: that's right. Absolutely no rhyme or reason.

I suppose the number naming convention is the logical corollary of a letter naming convention.
See my very first entry on the huge benefits to the Finnish of a sensible alphabet and spelling system.

Thursday, January 29, 2009

comparison of VERSE and LangX. peek at first HIOXian character.



D: I really must use a proper font creation program.
Sorry about the sketchy diagrams. Nonetheless, I think it shows the gist of HIOXian.
This particular character is the "TH" sound in 'loathe' and not not 'loath'.

You can now see the stylized anatomical figure in cross section that HIOXian represents.
In this case, the consonant is a square in the bottom 2/3 of the character space.
An optional diacritic can be placed above it.
Vowels would occupy the TOP 2/3 square of the space. Or v.v..
I think Roman alphabet conventions would make placing the vowels low more intuitive.

I now think I will simply not indicate anatomical parts that are disengaged, thus freeing up those bar segments for additional miscellaneous meaning. Cursive will also be faster.
I indicated an example of cursive writing. I confess I was inspired by Octomatic's approach.

Try to say "think". You likely did not say th-i-n-k. You likely said th-i-NG-k, even though that is NOT what the letters are. Biomechanics. It is easier. Your tongue is lazy. This is no surprise. The nerves send signals along routes of various lengths. Musculature varies, as does mass at rest and inertia of various body parts. It is a wonder we can be understood at all!
At any rate, now imagine (yes wave hands here until I finish it) the symbols for N and K side-by-side. Recall again how Octomatics allows arithmetic visually and without understanding of conventional mathematics. Now imagine a simple comparison of two bars, one in each HIOXian character. Using a simple comparison chart, somebody unable to use the IPA system would be able to point out that N is likely to deform into NG.
And so on.

I need to pay money for a decent font creation program. That means I need a credit card...

--------------------------
A comparison of LangX and VERSE. Time lines.

Provisional IAL Name

Number of Consonants and Vowels

Inaugural Year as Official IAL

First Language or Mother Tongue

Second or Auxiliary Language






Lang53

27 C 26 V

2726 AD

100%

0%






Lang49

26 C 23 V

2623 AD

98%

2%






Lang45

25 C 20 V

2520 AD

90%

10%






Lang41

24 C 17 V

2417 AD

70%

30%






Lang37

23 C 14 V

2314 AD

30%

70%






Lang33

22 C 11 V

2211 AD

10%

90%






Lang29

21 C 8 V

2108 AD

2%

98%






Lang25

20 C 5 V

2005 AD

0%

100%


D: again, I'd be concerned about AI/transhumanism on that time scale.

For the purpose of my sci-fi story, the musical pitch system of VERSE (VERy Simplified English) will be as follows.

1 Pitch 2020 AD
3 Pitch 2040 AD (me, fa, so)
7 pitch 2060 AD (do re mi fa so la ti)
12 Pitch 2080 AD (also half notes, Western-style)
20+ Pitch 2100 AD (also quarter notes, Eastern-style)

See my first blog entry for VERSE.
Basically, the hypothetical future scenario for an English Pidgin is as follows:
1) We start with Standard English
2) a mixing-pot culture of non-academics has a need for a useful fast new interlang
3) Verb affixes are removed and all such functions will be handled by particles.
4) Optionally, the particles are replaced with pitch on the simple basic verb.
5) More and more verb aspects are absorbed by more and more pitches.
This process is optional. Just like we talk down to children with simpler speech, so too could a future society replace pitch with the older verb modifier particles.
Optionally, BOTH could be used simultaneously and redundantly. This would be used for when a message MUST be very clear. It would likely also be used in a patronizing fashion.
In an emergency situation, when time is critical, only pitch would be used.

The example I used on the VERSE page was the following sentence.
"Dog bite boy." That is not proper English. So this must be early pidgin.
We all know how to modify the verb bite.
So too a noun, subject or direct/indirect object.
Bite, bit, biting, will be biting, et al.
A dog, some dogs, the dog, the(s) dogs, those/these, this/that, yadda yadda.

Language X/ Language 53 proposes reintroducing synthetic language elements gradually. However, my VERSE sci-fi scenario is for a subculture of submariner sea-faring nomads.
They have different demands for a language.
I saw a show on a crisis in a jumbo jet. I think the engine had failed or some such. Human factors analysis indicated that for brief periods English transmitted meaning at about 1 unit of meaning per second. I pondered how all those articles and auxiliary verbs and whatnot slowed them down. Of course, those grammatical elements were necessary for clear meaning, so were still used.

Well, affixes/particles (for synthetic and analytic respectively) are both essentially, to use an analogy, two-dimensional. They need to fit in their own part of time for communication.
They cannot be superimposed on top of root/stem words in a "three dimensional" sort of way.
Pitch essentially introduces a Cartesian Z co-ordinate. It takes the 'square' of words and makes them cubic, with depth.

VERSE is highly compatible with the LangX early stages. In fact, it can be seen as an alternative future route for LangX, once the basic international creole-esque vocabulary has been established and synthetic words have all been broken up into separate word particles.

Let us revisit that one sentence for a moment.
Dog bite boy.
There are a number of Eastern music octaves. Typically they have about 20 pitches, at quarter pitch-note intervals.
Three words.
Noun (subject) Verb Noun (direct object).
20x20x20.
20x20 is 400. 400x20 is 8000 permutations. In three words. In 3 syllables.
A language with optional agreement, with a simple mechanism to talk to early generations of speakers.
I need to hash out a similar system for prepositions. I am sure I can keep them down to one syllable also.

Aside: I was on a bus the other day. Looking at the 'next stop' sign, I was counting the pixels and studying the characters. They looked like Ygyde. OK, my objection about 45 degree angles was not valid. Those map just fine onto a low resolution monitor.