220 10559 <789a95ec-97de-4137-b28d-6be999867cbf@isocpp.org> article
Path: news.gmane.org!not-for-mail
From: Diggory Blake <diggsey@googlemail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Unicode support in the Standard Library
Date: Thu, 8 May 2014 18:40:30 -0700 (PDT)
Lines: 200
Approved: news@gmane.org
Message-ID: <789a95ec-97de-4137-b28d-6be999867cbf@isocpp.org>
References: <4ef82544-cd98-4488-8230-88ddaea78562@isocpp.org> <CAGNvRgA4keukGYJG_Z0OAiHeBnKa55K5X=4EYdtvHxq6tT4G4w@mail.gmail.com> <002D029C-6783-4A62-8CC4-B32B7BE8B23D@gmail.com>
 <4DA0EFEA-8C4E-499A-A8CC-37419B00D02E@gmail.com>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: multipart/alternative; 
	boundary="----=_Part_902_25792550.1399599630498"
X-Trace: ger.gmane.org 1399599639 14560 80.91.229.3 (9 May 2014 01:40:39 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Fri, 9 May 2014 01:40:39 +0000 (UTC)
To: std-proposals@isocpp.org
Original-X-From: std-proposals+bncBC2MLAWQ6ANBBD7EWCNQKGQES7MGAVA@isocpp.org Fri May 09 03:40:34 2014
Return-path: <std-proposals+bncBC2MLAWQ6ANBBD7EWCNQKGQES7MGAVA@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-pd0-f198.google.com ([209.85.192.198])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBC2MLAWQ6ANBBD7EWCNQKGQES7MGAVA@isocpp.org>)
	id 1WiZnd-0001hH-Ck
	for gclcip-std-proposals@m.gmane.org; Fri, 09 May 2014 03:40:33 +0200
Original-Received: by mail-pd0-f198.google.com with SMTP id w10sf13206448pde.5
        for <gclcip-std-proposals@m.gmane.org>; Thu, 08 May 2014 18:40:32 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=googlemail.com; s=20120113;
        h=date:from:to:message-id:in-reply-to:references:subject:mime-version
         :x-original-sender:reply-to:precedence:mailing-list:list-id
         :list-post:list-help:list-archive:list-subscribe:list-unsubscribe
         :content-type;
        bh=HJ759DL/gsk0HrB1HRMgSddh2pFsovYdfUz/LjbhN3c=;
        b=gha9+ZulHI6j3eExKs6bBVzZ9zwRE7mFZnlnSt83Pf7bWd8vRE+/BXp1idP7DXfUEK
         jcH7Si5kpuU4HSbniqYCVwEhRKTP2afqaqLc9xNVqL3cRDsomsV44nnFXbdvFfuxcfP6
         H7Rw/BHJPq1qx3WuKa/zKH95adwgGoTuPCHoB8mO3oYUAmvzD8ACG5fDfs10UbVIvUbS
         +zQ6uzmnZzsWB28ITYFUSaHKEpFq50rL2JjSzTyRsg9Z0sTQ9JjZ/dfEMl8YXW9Ko/Jp
         /SuJC8vC3icb7pjcfckiB4OIMDpiwDn1IOxKRt4l4mHlRKbmMFkF99wi2mzlqTcSF0S4
         dO2g==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:date:from:to:message-id:in-reply-to:references
         :subject:mime-version:x-original-sender:reply-to:precedence
         :mailing-list:list-id:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe:content-type;
        bh=HJ759DL/gsk0HrB1HRMgSddh2pFsovYdfUz/LjbhN3c=;
        b=VPbq7IXElt59OiylojLKV71g8gPJa5cuTNq2ozuWyMuz14pjOyt5m4epyxvPYk8yZv
         rRYSkK9vWvX0CaV8fKmy+c1UF6IAqwGmFXsxlx3ton62yvb1Faf7fU3Ky+DyGq8d8Up8
         ZhycIUxk6IyEFd6n0i89/whhuDjm2FOa4xMFvcMxlEoNueZE05gte1h4QSmTYZRmTVd6
         OINezc465CYmAIOxMP3D/XwkVqVrw1EB0Cg0I1MD8D5HDjBS6koJPRQ/B/44aeNbQcR9
         NBe5G5OL8eGLc+fw0+sD7pNDLyh7Ysf9Usmvz2/szKDRLSVQ9suUEbBckTrSvDwNGdol
         P5YA==
X-Gm-Message-State: ALoCoQndzfkwHMq8Ye8NTkLPSuxLnmnIEsgWOMfEokzXcxCl7jBlnvoSKqJlHMQSSPfWSDFXGQ8U
X-Received: by 10.66.230.226 with SMTP id tb2mr1113696pac.41.1399599632221;
        Thu, 08 May 2014 18:40:32 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.140.22.145 with SMTP id 17ls103334qgn.21.gmail; Thu, 08 May
 2014 18:40:31 -0700 (PDT)
X-Received: by 10.140.94.169 with SMTP id g38mr2126qge.13.1399599631325;
        Thu, 08 May 2014 18:40:31 -0700 (PDT)
In-Reply-To: <4DA0EFEA-8C4E-499A-A8CC-37419B00D02E@gmail.com>
X-Original-Sender: diggsey@googlemail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:10559
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/10559>

------=_Part_902_25792550.1399599630498
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable



On Friday, 9 May 2014 01:49:24 UTC+1, Dietmar K=C3=BChl wrote:
>
> OK, when I made the comment quoted below I was at work typing on a mobile=
=20
> device, hence, the message didn't contain all the parts which seem to be=
=20
> necessary to this discussion. I read all the other contributions and they=
=20
> are mostly going of into an area I consider the entirely wrong direction =
so=20
> I'll pretend they were not made and continue from this earlier point of t=
he=20
> discussion. So let me put the arguments for the general design together.=
=20
> Note, however, that I'm not going to write a proposal or make a promise t=
o=20
> review proposals made by other. However, the arguments below will guide m=
y=20
> arguments in the committee (unless someone makes good arguments that they=
=20
> are wrong.=20
>
> Step 1: External vs. Internal Encoding=20
>
> When processing strings there is always an encoding involved. In its=20
> simplest form, it is a singly byte, fixed width encoding like, e.g., ASCI=
I.=20
> It doesn't matter whether the characters are internal to a program, i.e.,=
=20
> they are stored in memory by the program or they are external to a progra=
m,=20
> i.e., they are in a file, in a buffer just read into a program, etc.: the=
=20
> is an encoding. However, it is important to realise that there is only=20
> *one* internal encoding and string shall be converted from whatever=20
> external encoding into the internal encoding upon reading and converted=
=20
> from the internal encoding to the external encoding upon writing! Dealing=
=20
> with multiple internal encodings [for the same character type] is neither=
=20
> necessary nor helpful. OK, it may be necessary if there are potential=20
> internal encodings which don't cover the same set of characters. Well, i=
=20
> was the case for single byte fixed width encodings: for example, the=20
> different choices of ISO-Latin-n covered different characters. However, w=
e=20
> are talking about Unicode processing and despite all its failures Unicode=
=20
> covers the full range of [human] characters (yes, Klingon characters were=
=20
> removed from Unicode; as far as I can tell to make space for a=20
> comprehensive set of characters for turds).=20
>

You're forgetting that C++ code may have to interoperate with other code=20
which uses a different internal encoding. For example, if I want to call=20
MessageBoxW on windows, I need to pass in a string which is utf16 encoded.=
=20
If I'm going to be calling such a method a lot I should be able to store=20
that string internally in utf16 so I don't have to convert every time.=20
Furthermore, the type system should prevent me from inadvertently assigning=
=20
non-utf16-encoded strings to it and vice-versa. This can be done without=20
modifying the basic_string class as per my previous suggestion. You're=20
right that 'char', 'char16_t' and 'char32_t' might not use the utf=20
encodings, so I would amend my suggestion such that=20
"char_traits<char/char16_t/char32_t>::encoding" would be the implementation=
=20
specific encoding used by string literals of the same type, rather than=20
necessarily utf8/16/32.


> > On 8 May 2014, at 15:30, Dietmar Kuehl <dietma...@gmail.com<javascript:=
>>=20
> wrote:=20
> > I will give the feedback I gave before: don't create another string=20
> class! Instead, create the necessary algorithms to deal with Unicode. In =
my=20
> opinion the actual encoding/decoding business is covered by the=20
> std::codecvt<...> facet although it may be worth explicitly defining=20
> instances of these facets for the various Unicode encodings. There are, o=
f=20
> course, plenty of other algorithms in Unicode which are reasonable to=20
> expose. Given that people like to process UTF8 and UTF16 it may be=20
> reasonable to also have encoding aware algorithms for string operations.=
=20
> >=20
> > Unless soneone provides a really strong argument for another string=20
> class, I will strongly argue against adding another representation for=20
> strings! (I can see a place for an immutable string class but that's=20
> entirely different).=20
>
>

--=20

---=20
You received this message because you are subscribed to the Google Groups "=
ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposa=
ls/.

------=_Part_902_25792550.1399599630498
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><br>On Friday, 9 May 2014 01:49:24 UTC+1, Dietmar K=C3=
=BChl  wrote:<blockquote class=3D"gmail_quote" style=3D"margin: 0;margin-le=
ft: 0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;">OK, when I made t=
he comment quoted below I was at work typing on a mobile device, hence, the=
 message didn't contain all the parts which seem to be necessary to this di=
scussion. I read all the other contributions and they are mostly going of i=
nto an area I consider the entirely wrong direction so I'll pretend they we=
re not made and continue from this earlier point of the discussion. So let =
me put the arguments for the general design together. Note, however, that I=
'm not going to write a proposal or make a promise to review proposals made=
 by other. However, the arguments below will guide my arguments in the comm=
ittee (unless someone makes good arguments that they are wrong.
<br>
<br>Step 1: External vs. Internal Encoding
<br>
<br>When processing strings there is always an encoding involved. In its si=
mplest form, it is a singly byte, fixed width encoding like, e.g., ASCII. I=
t doesn't matter whether the characters are internal to a program, i.e., th=
ey are stored in memory by the program or they are external to a program, i=
..e., they are in a file, in a buffer just read into a program, etc.: the is=
 an encoding. However, it is important to realise that there is only *one* =
internal encoding and string shall be converted from whatever external enco=
ding into the internal encoding upon reading and converted from the interna=
l encoding to the external encoding upon writing! Dealing with multiple int=
ernal encodings [for the same character type] is neither necessary nor help=
ful. OK, it may be necessary if there are potential internal encodings whic=
h don't cover the same set of characters. Well, i was the case for single b=
yte fixed width encodings: for example, the different choices of ISO-Latin-=
n covered different characters. However, we are talking about Unicode proce=
ssing and despite all its failures Unicode covers the full range of [human]=
 characters (yes, Klingon characters were removed from Unicode; as far as I=
 can tell to make space for a comprehensive set of characters for turds).
<br></blockquote><div><br>You're forgetting that C++ code may have to inter=
operate with other code which uses a different internal encoding. For examp=
le, if I want to call MessageBoxW on windows, I need to pass in a string wh=
ich is utf16 encoded. If I'm going to be calling such a method a lot I shou=
ld be able to store that string internally in utf16 so I don't have to conv=
ert every time. Furthermore, the type system should prevent me from inadver=
tently assigning non-utf16-encoded strings to it and vice-versa. This can b=
e done without modifying the basic_string class as per my previous suggesti=
on. You're right that 'char', 'char16_t' and 'char32_t' might not use the u=
tf encodings, so I would amend my suggestion such that "char_traits&lt;char=
/char16_t/char32_t&gt;::encoding" would be the implementation specific enco=
ding used by string literals of the same type, rather than necessarily utf8=
/16/32.<br><br></div><blockquote class=3D"gmail_quote" style=3D"margin: 0;m=
argin-left: 0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;">
<br>&gt; On 8 May 2014, at 15:30, Dietmar Kuehl &lt;<a href=3D"javascript:"=
 target=3D"_blank" gdf-obfuscated-mailto=3D"1vdOz7_9qJYJ" onmousedown=3D"th=
is.href=3D'javascript:';return true;" onclick=3D"this.href=3D'javascript:';=
return true;">dietma...@gmail.com</a>&gt; wrote:
<br>&gt; I will give the feedback I gave before: don't create another strin=
g class! Instead, create the necessary algorithms to deal with Unicode. In =
my opinion the actual encoding/decoding business is covered by the std::cod=
ecvt&lt;...&gt; facet although it may be worth explicitly defining instance=
s of these facets for the various Unicode encodings. There are, of course, =
plenty of other algorithms in Unicode which are reasonable to expose. Given=
 that people like to process UTF8 and UTF16 it may be reasonable to also ha=
ve encoding aware algorithms for string operations.
<br>&gt;=20
<br>&gt; Unless soneone provides a really strong argument for another strin=
g class, I will strongly argue against adding another representation for st=
rings! (I can see a place for an immutable string class but that's entirely=
 different).
<br>
<br></blockquote></div>

<p></p>

-- <br />
<br />
--- <br />
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br />
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to <a href=3D"mailto:std-proposals+unsubscribe@isocpp.org">std-proposa=
ls+unsubscribe@isocpp.org</a>.<br />
To post to this group, send email to <a href=3D"mailto:std-proposals@isocpp=
..org">std-proposals@isocpp.org</a>.<br />
Visit this group at <a href=3D"http://groups.google.com/a/isocpp.org/group/=
std-proposals/">http://groups.google.com/a/isocpp.org/group/std-proposals/<=
/a>.<br />

------=_Part_902_25792550.1399599630498--

.
