220 19855 <6e550a15-b9c7-4f95-88ca-725c500ae47a@isocpp.org> article
Path: news.gmane.org!not-for-mail
From: glen stark <g.a.stark@gmail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Unicode support in the Standard Library
Date: Fri, 14 Aug 2015 02:41:41 -0700 (PDT)
Lines: 93
Approved: news@gmane.org
Message-ID: <6e550a15-b9c7-4f95-88ca-725c500ae47a@isocpp.org>
References: <4ef82544-cd98-4488-8230-88ddaea78562@isocpp.org>
 <CAGNvRgA4keukGYJG_Z0OAiHeBnKa55K5X=4EYdtvHxq6tT4G4w@mail.gmail.com>
 <002D029C-6783-4A62-8CC4-B32B7BE8B23D@gmail.com>
 <CAGNvRgB5xSjojj2cZjBaaG=hTcBuRrXJR39xDv0QG1HuQLJ_gQ@mail.gmail.com>
 <01f2cc36-e743-4782-891e-c074dc072c7f@isocpp.org>
 <dac74199-e38a-47b8-b383-2623c34d6fe4@isocpp.org>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: multipart/mixed; 
	boundary="----=_Part_330_1802199779.1439545301435"
X-Trace: ger.gmane.org 1439545305 31643 80.91.229.3 (14 Aug 2015 09:41:45 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Fri, 14 Aug 2015 09:41:45 +0000 (UTC)
To: ISO C++ Standard - Future Proposals <std-proposals@isocpp.org>
Original-X-From: std-proposals+bncBDKMLD5P4YGBBVXPW2XAKGQEXXFEZPQ@isocpp.org Fri Aug 14 11:41:45 2015
Return-path: <std-proposals+bncBDKMLD5P4YGBBVXPW2XAKGQEXXFEZPQ@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-ig0-f197.google.com ([209.85.213.197])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBDKMLD5P4YGBBVXPW2XAKGQEXXFEZPQ@isocpp.org>)
	id 1ZQBUe-0000jx-I9
	for gclcip-std-proposals@m.gmane.org; Fri, 14 Aug 2015 11:41:44 +0200
Original-Received: by igcse8 with SMTP id se8sf19676549igc.1
        for <gclcip-std-proposals@m.gmane.org>; Fri, 14 Aug 2015 02:41:43 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=gmail.com; s=20120113;
        h=date:from:to:message-id:in-reply-to:references:subject:mime-version
         :content-type:x-original-sender:reply-to:precedence:mailing-list
         :list-id:x-spam-checked-in-group:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe;
        bh=PGEpYbzq7qFW6sHb6UKLepqBSRFebfwCB7LS1LCY11A=;
        b=Quxbr8ELtKZjorGMjPj6v5hBIo0LXLj50GiAK/CORZ9iCgbW3022Z3Unng6vC4GZSJ
         SgcxqbXyIFEqQ+Egox/i/Vva0Zx8JpJktEw1e9SFjdswlHYEQlNRibLe7s4YuiarIC4Z
         fs/YF/Ch5NYSckJqrPxkrcD15RC3R2PkR7eymVNbrOHl1nkeGmYLYP4S/MqzjgI3LRFq
         59ZpYJKHVfrAsEykUcWzulvW7q6yOtbCsP1FN22gzRGdS8yugUMGqyKHUEy1QEmzNyA+
         YNF201rzJnHBKuPcbBfFLIlzOISMMZ5onzJZ+5iXJG5aYWmJo1j9L+Cfmx8XyV7qKJ7w
         xkLg==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:date:from:to:message-id:in-reply-to:references
         :subject:mime-version:content-type:x-original-sender:reply-to
         :precedence:mailing-list:list-id:x-spam-checked-in-group:list-post
         :list-help:list-archive:list-subscribe:list-unsubscribe;
        bh=PGEpYbzq7qFW6sHb6UKLepqBSRFebfwCB7LS1LCY11A=;
        b=aNO6QU39QKFEoMw8s9HLxlOabgexeAjyBhkslZOL1+iT93RGEXUPVcU//dysn/vkPR
         RSzV5g8hhvDq9XlNtfKvEugPs1P/itW9Fob9NSoXLu7F5OlhaDlR39iJqx/IvWSd4RZ9
         MeAcIrktwjabiuTgImvmSyjc4qFnHCMrdOeroYyyFb+kmTiua63jtjlnjcMa0J41rd8W
         8gtdKQUkfO6XuRIlb8NPH9OqrSAqVEFulOWbigG1iK3vQ9JK2RnwQpBlpEKi8rfLTk6F
         qPvcrJ/5TuXnL0MtYz4p0/WUjO9vWaUjKxm1CQ2OXi+KEsM6rtxLm2n/4SkEFn4/sRmT
         iCsw==
X-Gm-Message-State: ALoCoQmuHyUnIPqXxcNpBYPXMOXcOp+opbYZ1D2wu098yMe2QuqTH//V8zbaG8lkWb+nAJrlmLXP
X-Received: by 10.107.15.18 with SMTP id x18mr40502658ioi.28.1439545303355;
        Fri, 14 Aug 2015 02:41:43 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.140.88.233 with SMTP id t96ls1260509qgd.78.gmail; Fri, 14 Aug
 2015 02:41:42 -0700 (PDT)
X-Received: by 10.140.105.194 with SMTP id c60mr330213qgf.36.1439545302066;
        Fri, 14 Aug 2015 02:41:42 -0700 (PDT)
In-Reply-To: <dac74199-e38a-47b8-b383-2623c34d6fe4@isocpp.org>
X-Original-Sender: g.a.stark@gmail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Spam-Checked-In-Group: std-proposals@isocpp.org
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>,
 <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:19855
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/19855>

------=_Part_330_1802199779.1439545301435
Content-Type: multipart/alternative; 
	boundary="----=_Part_331_1083330414.1439545301435"

------=_Part_331_1083330414.1439545301435
Content-Type: text/plain; charset=UTF-8



On Thursday, May 8, 2014 at 5:09:10 PM UTC+2, Diggory Blake wrote:
>
> The thing is, for most of the time the encoding of the string is not 
> relevant - copying, appending, exact comparison all work on the raw data 
> and don't care about the encoding. The only time the encoding matters is 
> during I/O or when performing lexicographic operations, and generally 
> within the same application, the same encoding will be used (almost) 
> everywhere.
>
>>
>>>
While it's probably true that within most applications, the same encoding 
will be used (almost) everywhere,  that wouldn't be true for libraries. 
 For library support, and for unicode-aware components of legacy 
applications, I think it would be helpful to know, if you're getting a 
std::basic_string<char>, if that's ASCII, latin1, or utf-8.    

Would it be reasonable to add encoding information to char_traits?  This 
would allow a library developer that has to support both legacy code and 
utf-8 compliant code to use type information to ensure correct behavior, 
and prevent accidental operations which are meaningless across encodings. 
 If it is reasonable, it would take some thought:  ideally the standard 
explicitly provide encodings that are likely to remain in use for the next 
decade or two (latin1  comes to mind), and it should be possible to create 
custom encodings.  

-- 

--- 
You received this message because you are subscribed to the Google Groups "ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an email to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposals/.

------=_Part_331_1083330414.1439545301435
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><br>On Thursday, May 8, 2014 at 5:09:10 PM UTC+2, Digg=
ory Blake wrote:<blockquote class=3D"gmail_quote" style=3D"margin: 0;margin=
-left: 0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;"><div dir=3D"lt=
r">The thing is, for most of the time the encoding of the string is not rel=
evant - copying, appending, exact comparison all work on the raw data and d=
on&#39;t care about the encoding. The only time the encoding matters is dur=
ing I/O or when performing lexicographic operations, and generally within t=
he same application, the same encoding will be used (almost) everywhere.<br=
><blockquote class=3D"gmail_quote" style=3D"margin:0;margin-left:0.8ex;bord=
er-left:1px #ccc solid;padding-left:1ex"><div dir=3D"ltr"><div><blockquote =
class=3D"gmail_quote" style=3D"margin:0;margin-left:0.8ex;border-left:1px #=
ccc solid;padding-left:1ex"><br></blockquote></div></div></blockquote></div=
></blockquote><div><br></div><div>While it&#39;s probably true that within =
most applications, the same encoding will be used (almost) everywhere, =C2=
=A0that wouldn&#39;t be true for libraries. =C2=A0For library support, and =
for unicode-aware components of legacy applications, I think it would be he=
lpful to know, if you&#39;re getting a std::basic_string&lt;char&gt;, if th=
at&#39;s ASCII, latin1, or utf-8. =C2=A0 =C2=A0</div><div><br></div><div>Wo=
uld it be reasonable to add encoding information to char_traits? =C2=A0This=
 would allow a library developer that has to support both legacy code and u=
tf-8 compliant code to use type information to ensure correct behavior, and=
 prevent accidental operations which are meaningless across encodings. =C2=
=A0If it is reasonable, it would take some thought: =C2=A0ideally the stand=
ard explicitly provide encodings that are likely to remain in use for the n=
ext decade or two (latin1 =C2=A0comes to mind), and it should be possible t=
o create custom encodings. =C2=A0</div></div>

<p></p>

-- <br />
<br />
--- <br />
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br />
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to <a href=3D"mailto:std-proposals+unsubscribe@isocpp.org">std-proposa=
ls+unsubscribe@isocpp.org</a>.<br />
To post to this group, send email to <a href=3D"mailto:std-proposals@isocpp=
..org">std-proposals@isocpp.org</a>.<br />
Visit this group at <a href=3D"http://groups.google.com/a/isocpp.org/group/=
std-proposals/">http://groups.google.com/a/isocpp.org/group/std-proposals/<=
/a>.<br />

------=_Part_331_1083330414.1439545301435--
------=_Part_330_1802199779.1439545301435--

.
