220 10558 <4DA0EFEA-8C4E-499A-A8CC-37419B00D02E@gmail.com> article
Path: news.gmane.org!not-for-mail
From: Dietmar Kuehl <dietmar.kuehl@gmail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Unicode support in the Standard Library
Date: Fri, 9 May 2014 01:49:24 +0100
Lines: 203
Approved: news@gmane.org
Message-ID: <4DA0EFEA-8C4E-499A-A8CC-37419B00D02E@gmail.com>
References: <4ef82544-cd98-4488-8230-88ddaea78562@isocpp.org> <CAGNvRgA4keukGYJG_Z0OAiHeBnKa55K5X=4EYdtvHxq6tT4G4w@mail.gmail.com> <002D029C-6783-4A62-8CC4-B32B7BE8B23D@gmail.com>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0 (Mac OS X Mail 7.2 \(1874\))
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable
X-Trace: ger.gmane.org 1399596579 9841 80.91.229.3 (9 May 2014 00:49:39 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Fri, 9 May 2014 00:49:39 +0000 (UTC)
To: "std-proposals@isocpp.org" <std-proposals@isocpp.org>
Original-X-From: std-proposals+bncBDZYDW6QHYIJTTFQTMCRUBA7Y6SSU@isocpp.org Fri May 09 02:49:31 2014
Return-path: <std-proposals+bncBDZYDW6QHYIJTTFQTMCRUBA7Y6SSU@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-ee0-f71.google.com ([74.125.83.71])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBDZYDW6QHYIJTTFQTMCRUBA7Y6SSU@isocpp.org>)
	id 1WiZ0E-0000jY-W4
	for gclcip-std-proposals@m.gmane.org; Fri, 09 May 2014 02:49:31 +0200
Original-Received: by mail-ee0-f71.google.com with SMTP id c13sf2118481eek.10
        for <gclcip-std-proposals@m.gmane.org>; Thu, 08 May 2014 17:49:30 -0700 (PDT)
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:to:references:in-reply-to:mime-version
         :message-id:from:subject:date:x-original-sender
         :x-original-authentication-results:reply-to:precedence:mailing-list
         :list-id:list-post:list-help:list-archive:list-subscribe
         :list-unsubscribe:content-type:content-transfer-encoding;
        bh=6RVBqW/Y5mokIaSOEKsiyZS8yOF8TM+J0ReChUoLSoA=;
        b=Wv+TgAF7mQjVImzg2QdvgMuR85aAUyAfPKhAZ0snujQwsq7UnxnXPt+edMVhsffTKg
         0Yb2+av67IdlHcTZwb1K41zzwFipEFGFtZeH3U+MqAJcZAZ2d4i0xuQ+px4JJ8F9ZuuM
         e1gA7ec2bmGm+whjbEOrcumncdRRav+gm875ph08bzbs3WRkce57LVKmBgTv3zLFnLE/
         yh+cX+5duBWgoAGS3kJaWYfwNyo4TOV0zH0sAHf9X1i8gF91SpcFkwrj7wWMQzSwOJ3B
         UeahsBYiPYWefXe1lOX9JHcL2LYuoBrMT2w0OfsVfWoWGY3CgC1MNHlge5uszrHsNLM0
         OEzw==
X-Gm-Message-State: ALoCoQnVDsMezTSToCs8XxiUBSlsp8Nikl4fOkpoinNRHjutTE1Fb0a05g6i1tnKGCdyEaaISjwW
X-Received: by 10.112.184.227 with SMTP id ex3mr17222lbc.16.1399596570467;
        Thu, 08 May 2014 17:49:30 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.180.206.67 with SMTP id lm3ls48242wic.1.gmail; Thu, 08 May
 2014 17:49:28 -0700 (PDT)
X-Received: by 10.194.80.7 with SMTP id n7mr5782052wjx.8.1399596568956;
        Thu, 08 May 2014 17:49:28 -0700 (PDT)
Original-Received: from mail-wi0-x22c.google.com (mail-wi0-x22c.google.com [2a00:1450:400c:c05::22c])
        by mx.google.com with ESMTPS id c2si119229wjf.29.2014.05.08.17.49.28
        for <std-proposals@isocpp.org>
        (version=TLSv1 cipher=ECDHE-RSA-RC4-SHA bits=128/128);
        Thu, 08 May 2014 17:49:28 -0700 (PDT)
Received-SPF: pass (google.com: domain of dietmar.kuehl@gmail.com designates 2a00:1450:400c:c05::22c as permitted sender) client-ip=2a00:1450:400c:c05::22c;
Original-Received: by mail-wi0-f172.google.com with SMTP id hi2so552677wib.11
        for <std-proposals@isocpp.org>; Thu, 08 May 2014 17:49:28 -0700 (PDT)
X-Received: by 10.180.82.7 with SMTP id e7mr787511wiy.6.1399596568644;
        Thu, 08 May 2014 17:49:28 -0700 (PDT)
Original-Received: from [192.168.0.7] (cpc18-dals16-2-0-cust21.20-2.cable.virginm.net. [81.100.53.22])
        by mx.google.com with ESMTPSA id mw4sm2122190wib.12.2014.05.08.17.49.27
        for <std-proposals@isocpp.org>
        (version=TLSv1 cipher=ECDHE-RSA-RC4-SHA bits=128/128);
        Thu, 08 May 2014 17:49:27 -0700 (PDT)
X-Universally-Unique-Identifier: 9CE9F0F2-A367-4F47-BDE6-1F6C8D33DC70
In-Reply-To: <002D029C-6783-4A62-8CC4-B32B7BE8B23D@gmail.com>
X-Smtp-Server: smtp.me.com:dietmar.kuehl
X-Apple-Mail-Signature: 
X-Mailer: iPhone Mail (11D201)
X-Apple-Windows-Friendly: 1
X-Apple-Base-Url: x-msg://8/
X-Apple-Mail-Remote-Attachments: YES
X-Uniform-Type-Identifier: com.apple.mail-draft
X-Apple-Mail-Plain-Text-Draft: yes
X-Original-Sender: dietmar.kuehl@gmail.com
X-Original-Authentication-Results: mx.google.com;       spf=pass (google.com:
 domain of dietmar.kuehl@gmail.com designates 2a00:1450:400c:c05::22c as
 permitted sender) smtp.mail=dietmar.kuehl@gmail.com;       dkim=pass
 header.i=@gmail.com;       dmarc=pass (p=NONE dis=NONE) header.from=gmail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:10558
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/10558>

OK, when I made the comment quoted below I was at work typing on a mobile d=
evice, hence, the message didn't contain all the parts which seem to be nec=
essary to this discussion. I read all the other contributions and they are =
mostly going of into an area I consider the entirely wrong direction so I'l=
l pretend they were not made and continue from this earlier point of the di=
scussion. So let me put the arguments for the general design together. Note=
, however, that I'm not going to write a proposal or make a promise to revi=
ew proposals made by other. However, the arguments below will guide my argu=
ments in the committee (unless someone makes good arguments that they are w=
rong.

Step 1: External vs. Internal Encoding

When processing strings there is always an encoding involved. In its simple=
st form, it is a singly byte, fixed width encoding like, e.g., ASCII. It do=
esn't matter whether the characters are internal to a program, i.e., they a=
re stored in memory by the program or they are external to a program, i.e.,=
 they are in a file, in a buffer just read into a program, etc.: the is an =
encoding. However, it is important to realise that there is only *one* inte=
rnal encoding and string shall be converted from whatever external encoding=
 into the internal encoding upon reading and converted from the internal en=
coding to the external encoding upon writing! Dealing with multiple interna=
l encodings [for the same character type] is neither necessary nor helpful.=
 OK, it may be necessary if there are potential internal encodings which do=
n't cover the same set of characters. Well, i was the case for single byte =
fixed width encodings: for example, the different choices of ISO-Latin-n co=
vered different characters. However, we are talking about Unicode processin=
g and despite all its failures Unicode covers the full range of [human] cha=
racters (yes, Klingon characters were removed from Unicode; as far as I can=
 tell to make space for a comprehensive set of characters for turds).

With C++ there is a slight and somewhat annoying complication in that C++ h=
as multiple character types with different width. That is, different charac=
ter types will use different internal encodings. Since the strings will hav=
e different types there isn't much danger of accidental interference althou=
gh there is some danger as individual character types happily convert betwe=
en each other. It is worth to note that in the context of Unicode processin=
g the different *character* values actually do **not** represent *character=
s*! For most characters types one value cannot represent a character as the=
re aren't enough bits in the value (with char32_t being the exception; well=
, at least, for now). Instead, a character value (e.g. an individual `char`=
 or `wchar_t`) actually represents only a part of a character. Thinking mos=
tly in terms of strings of `char` representing something like UTF8 sequence=
, I think of the values stored in the character types as *bytes* although a=
 `wchar_t` is, at least, made up of two bytes (ICU calls what I call bytes =
*singletons*). That is, strings of the fundamental character types are made=
 up of bytes where one or more bytes form an actual character. Actually, mu=
ltiple bytes make up a *code point* which is used to represent a character =
but I will character and code point interchangeably.

Since C++ defines string literals for all of these character types (right n=
ow I'm not sure if there are really `signed char` and `unsigned char` liter=
als but I will simply assume that these will use the same encoding as `char=
`) and these string literals may contain Unicode characters, the internal e=
ncoding for each character types is actually already chosen by the compiler=
 (I think, the compilers are free to choose the encoding, though). That is,=
 **all** internal processing of Unicode strings will use the character type=
s specific internal encoding! There is no need to think about different enc=
odings [for a specific character type] internal to the program. This invari=
ant is extremely important and makes thinking about string processing possi=
ble. It also means, we can ignore the concept of encoding entirely for the =
internal processing of characters! This invariant is somewhat similar to th=
e approach in physics to choose the dimensions such that all the fundamenta=
l constants, e.g., the speed of light in vacuum (c), become 1: the resultin=
g math for the complicated formulas becomes viable (still incomprehensible =
to me but they look a huge amount simpler).

Why do encoding still matter? Well, outside of the program there are arbitr=
ary and often fairly odd encodings entirely outside the control of the prog=
ram (although there is a hope that in the long-term there will be just a fe=
w more or less reasonable encodings left). However, dealing with these exte=
rnal encodings can be centralised. In fact, dealing with external encodings=
 already *is* centralised: this is what the `std::codecvt<InternT, ExternT,=
 StateT>` facets are form. That is possibly not with an ideal interface and=
 we *may* consider creating a different interface but the conversion to and=
 from an external to encoding to the internal encoding is entirely orthogon=
al to a library processing Unicode.

Conveniently, the above means that for a Unicode library we can *entirely* =
ignore any encoding issues except, of course, the fact that the bytes used =
in strings are representing characters according to some encoding. What tha=
t encoding is, the programmer doesn't need to care about as the Unicode alg=
orithms will correctly interpret the bytes to process characters.

I have never used ICU directly but from a cursory look at their interfaces =
I *think* the ICU design makes the same assumption: once strings are inside=
 a program, their encoding is known. Actually, ICU further simplifies its v=
iew of the world by having just one character type, not 6 (`char`, `unsigne=
d char`, `signed char`, `wchar_t`, `char16_t`, and `char32_t`). I don't thi=
nk C++ has this luxury.

Step 2: Vocabulary Types

With encodings out of the way, let's now turn the focus on string types: ho=
w may should we have? Since strings are a fundamental abstraction which is =
used in many interfaces which need to communicate between different compone=
nts, the obvious answer is "one!" If there is more than one, some component=
s will choose to use different string types than other components. Sadly, s=
tring literals and the standard library string class have different types m=
eaning that there are already two string types - for each of the 6 choices =
of character types! While it would be borderline viable to unify the string=
 literals and `std::basic_string<cT>` using a suitable `std::basic_string_v=
iew<cT>`, a similar approach isn't possible for the different character typ=
es. To make matters worse, there are at least two popular choices for the c=
ommonly used character types (`wchar_t` on Windows and `char` everywhere el=
se). This situation is sufficiently bad and we should refrain from worsenin=
g the problem by introducing another string type: strings are fundamental v=
ocabulary and choice on the fundamental vocabulary is inherently bad.

To some extend there may be room for a different decision: instead of suppo=
rting all character types, it may be reasonable to support exactly one char=
acter type for the purpose of a Unicode library. Whenever an interface asks=
 for or provides a string of a different type, the strings would need to be=
 transformed which probably involves a change of the used encoding. Having =
just one string type to deal with is a huge advantage and, as far as I can =
tell, everybody always uses just one string type (either `std::string` or `=
std::wstring`). Since the standard would make a choice, the choice could be=
 deliberately to use `std::ustring` which could be mandated to be neither `=
std::string` nor `std::wstring` and drawing attention to the fact that `std=
::ustring` is guaranteed to use a Unicode encoding, i.e., the individual by=
tes (singletons) inside a `std::ustring` do **not** represent a character b=
ut merely part of a character. This assumption is often made for `std::stri=
ng` and `std::wstring`, too, but there is a lot of code out there which tre=
ats `std::string` and `std::wstring` as if they contained characters rather=
 than bytes/singletons.

Although the use of just one string class, i.e., a particular instantiation=
 of `std::basic_string<cT>` with a suitable character type `cT`, sounds gre=
at, I doubt that it will get enough support: either the `std::string` world=
 will be upset or the `std::wstring` or both. I honestly believe that we wo=
uld do ourselves and future generations of programmers a **tremendous** fav=
our if we could agree on The One string type (which might even use a polymo=
rphic allocator to remove the potential desire to vary allocation policies =
by changing the type; I'm personally on the fence on that one but that is i=
tself a bigger discussion which I'm not gone have now). Yes, I realise that=
 using just one string type would be disruptive now instead of taking out a=
 huge credit on the future we would remove that technical debt.

Step 3: How would a Unicode library look like?

First of all, even if there is no encoding to be dealt with and no new stri=
ng class, there are plenty of operations which are non-trivial, partly due =
to the way Unicode is designed:

- character-aware string processing: locating, extracting, changing, etc. i=
ndividual characters
- determining the number of characters, splitting strings after a certain n=
umber of characters,=20
- comparing strings: aside from Unicode strings not necessarily being norma=
lised, ordering strings in a form usable for humans is non-trivial: even tr=
ivial representation like the European letter-based strings are ordered dif=
ferently depending on the language context, e.g., where `ll` goes in the Sp=
anish or other European languages. This is nothing compared to ordering str=
ings with Chinese or Japanese characters.
- have a look at the operations in ICU to get more topics to be addressed.

The next question then becomes on how these algorithms should deal with str=
ings? The most likely abstraction is to have the algorithms operate in term=
s of character iterators, possibly limiting the set of iterator types the a=
lgorithms would operate on. I'm not sure if that is needed but I could imag=
ine that it may be undesirable to have all these algorithms be templates de=
fined in headers and instantiated wherever needed.

However, just processing STL-like iterators is probably not enough because =
sequences of characters may change the number of bytes used to encode them =
when applying even trivial transformations! That is, these algorithms may t=
ravel in terms of a richer abstraction than STL iterators. I haven't tried =
to create the corresponding abstractions.

This is what I think is important in this domain right from the top of my h=
ead. There is probably more but for now, I think the summary is:

- deal with only one encoding [per character type]
- do not create another string type; if absolutely necessary to have a type=
, create something viewing an underlying string type
- the algorithms operating on Unicode string are the important aspect

> On 8 May 2014, at 15:30, Dietmar Kuehl <dietmar.kuehl@gmail.com> wrote:
> I will give the feedback I gave before: don't create another string class=
! Instead, create the necessary algorithms to deal with Unicode. In my opin=
ion the actual encoding/decoding business is covered by the std::codecvt<..=
..> facet although it may be worth explicitly defining instances of these fa=
cets for the various Unicode encodings. There are, of course, plenty of oth=
er algorithms in Unicode which are reasonable to expose. Given that people =
like to process UTF8 and UTF16 it may be reasonable to also have encoding a=
ware algorithms for string operations.
>=20
> Unless soneone provides a really strong argument for another string class=
, I will strongly argue against adding another representation for strings! =
(I can see a place for an immutable string class but that's entirely differ=
ent).

--=20

---=20
You received this message because you are subscribed to the Google Groups "=
ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposa=
ls/.

.
