{"id":7083,"date":"2015-05-31T01:25:58","date_gmt":"2015-05-31T09:25:58","guid":{"rendered":"http:\/\/www.oeconomist.com\/blogs\/daniel\/?p=7083"},"modified":"2015-08-25T23:17:25","modified_gmt":"2015-08-26T07:17:25","slug":"a-question-of-characters","status":"publish","type":"post","link":"https:\/\/www.oeconomist.com\/blogs\/daniel\/?p=7083","title":{"rendered":"A Question of Characters"},"content":{"rendered":"<p>At various times, I'm confronted with confusion by persons and by systems of <span style=\"font-style: italic ;\">characters<\/span> with <span style=\"font-style: italic ;\">glyphs<\/span>.  Most of the time, that confusion is a very minor annoyance; sometimes, as when wrestling with the preparation of a technical document, it can cause many hours of difficulty.<\/p> <p>It's probably rather easier for people first to see that a <span style=\"font-style: italic ;\">character<\/span> may have multiple <span style=\"font-style: italic ;\">glyphs<\/span>.  For example, here are two distinct yet common <span style=\"font-style: italic ;\">glyphs<\/span> for the lower-case letter <q>a<\/q>: <image src=\"http:\/\/www.oeconomist.com\/blogs\/daniel\/wp-content\/uploads\/2015\/05\/char_glyph_01.png\" alt=\"\" width=\"110\" height=\"50\" style=\"display: block ; margin-left: auto ; margin-right: auto ; margin-top: 0.5em ; margin-bottom: 0.5em ; border: none ;\"><\/image> and here are two for <q>g<\/q>: <image src=\"http:\/\/www.oeconomist.com\/blogs\/daniel\/wp-content\/uploads\/2015\/05\/char_glyph_02.png\" alt=\"\" width=\"109\" height=\"53\" style=\"display: block ; margin-left: auto ; margin-right: auto ; margin-top: 0.5em ; margin-bottom: 0.5em ; border: none ;\"><\/image><\/p> <p>People have a bit more trouble with the idea that a single <span style=\"font-style: italic ;\">glyph<\/span> can correspond to more than one <span style=\"font-style: italic ;\">character<\/span>.  Perhaps most educated folk generally understand that a Greek <q>&Rho;<\/q> is <em>not<\/em> our <q>P<\/q>, even though one could easily imagine an identical <span style=\"font-style: italic ;\">glyph<\/span> being used in some fonts.  But many people think that they're looking at a <q>o<\/q> with an <span style=\"font-style: italic ;\">umlaut<\/span> in each of these two words: <image src=\"http:\/\/www.oeconomist.com\/blogs\/daniel\/wp-content\/uploads\/2015\/05\/char_glyph_03.png\" alt=\"\" width=\"303\" height=\"125\" style=\"display: block ; margin-left: auto ; margin-right: auto ; margin-top: 0.5em ; margin-bottom: 0.5em ; border: none ;\"><\/image> where&auml;s the two dots over the <q>o<\/q> in the first word are a <span style=\"font-style: italic ;\">di&aelig;resis<\/span>, an ancient diacritical mark used in various languages to clarify whether and how a vowel is pronounced.<span style=\"vertical-align: top ; font-size: smaller ;\">&#91;1&#93;<\/span>  The two dots over the <q>o<\/q> in the German <q><span style=\"font-style: italic ;\">sh&ouml;n<\/span><\/q> are indeed an <span style=\"font-style: italic ;\">umlaut<\/span>, which evolved far more recently from a <em>superscript <q>e<\/q><\/em>.<span style=\"vertical-align: top ; font-size: smaller ;\">&#91;2&#93;<\/span> (One may alternately write the same word <q><span style=\"font-style: italic ;\">schoen<\/span><\/q>, where&auml;s <q><span style=\"font-style: italic ;\">schon<\/span><\/q> is a different word.)<\/p> <p>Out of context, what one <em>sees<\/em> is a <span style=\"font-style: italic ;\">glyph<\/span>.  Generally, we need <em>context<\/em> to tell use whether we're looking at <q>&#1017;<\/q> (upper-case lunate sigma), our familiar <q>C<\/q>, or <q>&#1057;<\/q> (upper-case Cyrillic ess); likewise for many other <span style=\"font-style: italic ;\">characters<\/span> and their similar or identical <span style=\"font-style: italic ;\">glyphs<\/span>.  Until comparatively recently, we usually had sufficient context, mistakes were relatively infrequent and usually unimportant. (Okay, so a bunch of people thought that the Soviet Union called itself the <q>CCCP<\/q>, rather than the <q>&#1057;&#1057;&#1057;&#1056;<\/q>.  Meh.) But, with the development of electronic information technology, and with globalization, the distinction becomes more pressing.  Most of us have seen the problems of <abbr class=\"noshrink\" title=\"optical character recognition\">OCR<\/abbr>; these are essentially problems of inferring <span style=\"font-style: italic ;\">characters<\/span> from <span style=\"font-style: italic ;\">glyphs<\/span>.  It's not so messy when converting instead from plain-text or from something such as <abbr class=\"noshrink\" title=\"OpenDocument Format\">ODF<\/abbr>, but when <span style=\"font-style: italic ;\">character<\/span> substitutions were made based upon similarity or identity of <span style=\"font-style: italic ;\">glyph<\/span>, the <em>very<\/em> same problems can then arise.  For example, as I said, one <em>sees<\/em> <span style=\"font-style: italic ;\">glyphs<\/span>, but what is <em>heard<\/em> when the text is rendered audible will be phonetic values associated with the <span style=\"font-style: italic ;\">characters<\/span> used.  And sometimes the system will process a <span style=\"font-style: italic ;\">less-than<\/span> sign as a <span style=\"font-style: italic ;\">left angle bracket<\/span>, because everyone else is using it as such.  In an abstract sense, these are of course problems of <em>transliteration<\/em>, and of its effects upon <em>translation<\/em>.<\/p> <p>Some of you will recognize the contrast between <span style=\"font-style: italic ;\">character<\/span> and <span style=\"font-style: italic ;\">glyph<\/span> as a special case of the contrast between <span style=\"font-style: italic ;\">content<\/span> and <span style=\"font-style: italic ;\">presentation<\/span> &mdash; between <em>what one seeks to deliver<\/em> and <em>the manner of delivery<\/em>.  Some will also note that the boundary between the two shifts.  For example, the difference between upper-case and lower-case letters originated as nothing more than a difference in <span style=\"font-style: italic ;\">glyphs<\/span>.  Indeed, our <q>R<\/q> was once no more than a different way of writing the Greek <q>&Rho;<\/q>; our <q>A<\/q> simply <em>was<\/em> the Greek <q>&Alpha;<\/q>, and it can remain hard to distinguish them!  I don't know that <q>&#383;<\/q> (long ess) should be regarded as a different <em>character<\/em> from <q>s<\/q>, rather than just as an archa&iuml;c <span style=\"font-style: italic ;\">glyph<\/span> thereof.<\/p> <p>Still, the fact that what is sometimes mere <span style=\"font-style: italic ;\">presentation<\/span> may at other times <em>be<\/em> <span style=\"font-style: italic ;\">content<\/span> doesn't mean that we should forgo the gains to be had in being mindful of the distinction and in creating structures that <em>often<\/em> help us to avoid being shackled to the <span style=\"font-style: italic ;\">accidental<\/span>.<\/p> <hr width=\"50%\" align=\"left\" \/> <p><span style=\"vertical-align: top ; font-size: smaller ;\">&#91;1&#93;<\/span> In English and most other languages, a di&aelig;resis over the second of two vowels indicates that the vowel is pronounced separately, rather than forming a diphthong. (So here <span style=\"white-space: nowrap ;\">\/ko&#712;ap&#601;&#716;ret\/<\/span> rather than <span style=\"white-space: nowrap ;\">\/&#712;kup&#601;&#716;ret\/<\/span> or <span style=\"white-space: nowrap ;\">\/&#712;k&#650;p&#601;&#716;ret\/<\/span>.) Over a vowel standing alone, as in <q>Bront&euml;<\/q>, the di&aelig;resis signals that the vowel is not silent. (In English and some other languages, a grave accent may be used to the very same effect.) Portuguese cleverly uses a di&aelig;resis over the <em>first<\/em> of two vowels to signal that diphthong is formed where it might not be expected.<\/p> <p><span style=\"vertical-align: top ; font-size: smaller ;\">&#91;2&#93;<\/span> Germans used to use a dreadful script &mdash; <span style=\"font-style: italic ;\">Kurrentschrift<\/span> &mdash; in which such an evolution is less surprising.<\/p>","protected":false},"excerpt":{"rendered":"At various times, I'm confronted with confusion by persons and by systems of characters with glyphs. Most of the time, that confusion is a very minor annoyance; sometimes, as when wrestling with the preparation of a technical document, it can cause many hours of difficulty. It's probably rather easier for people first to see that [&hellip;]","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_bbp_topic_count":0,"_bbp_reply_count":0,"_bbp_total_topic_count":0,"_bbp_total_reply_count":0,"_bbp_voice_count":0,"_bbp_anonymous_reply_count":0,"_bbp_topic_count_hidden":0,"_bbp_reply_count_hidden":0,"_bbp_forum_subforum_count":0,"footnotes":""},"categories":[6,117,69,4],"tags":[793,1271,1273,1275,1270,1272,810,1274],"class_list":["post-7083","post","type-post","status-publish","format-standard","hentry","category-commentary","category-communication","category-information-technology","category-public","tag-characters","tag-content","tag-diaereses","tag-diereses","tag-glyphs","tag-presentation","tag-symbols","tag-umlauts"],"_links":{"self":[{"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=\/wp\/v2\/posts\/7083","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=7083"}],"version-history":[{"count":0,"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=\/wp\/v2\/posts\/7083\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=7083"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=7083"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.oeconomist.com\/blogs\/daniel\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=7083"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}