220 10647 <12556884-b2fa-4e85-b3b4-985defe7614e@isocpp.org> article
Path: news.gmane.org!not-for-mail
From: Diggory Blake <diggsey@googlemail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Unicode support in the Standard Library
Date: Thu, 15 May 2014 12:02:40 -0700 (PDT)
Lines: 164
Approved: news@gmane.org
Message-ID: <12556884-b2fa-4e85-b3b4-985defe7614e@isocpp.org>
References: <4ef82544-cd98-4488-8230-88ddaea78562@isocpp.org>
 <CAGNvRgA4keukGYJG_Z0OAiHeBnKa55K5X=4EYdtvHxq6tT4G4w@mail.gmail.com>
 <002D029C-6783-4A62-8CC4-B32B7BE8B23D@gmail.com> <CAGNvRgB5xSjojj2cZjBaaG=hTcBuRrXJR39xDv0QG1HuQLJ_gQ@mail.gmail.com>
 <01f2cc36-e743-4782-891e-c074dc072c7f@isocpp.org> <CAGNvRgBQeFOzzqgbWGMDrVEgmeHPPsG8FF0Q_M=ZyH52DbwsqA@mail.gmail.com>
 <1d79df40-3313-4241-a9b8-d9364d68d327@isocpp.org> <CAFk2RUb8CrW1NbLJLcfSvZX4XvuPoqTP9XwVDeC43-2YCQ5yrg@mail.gmail.com>
 <d7a9ccd4-8758-415e-918f-0865618ab07d@isocpp.org> <b59b8b8b-3073-4b02-892c-0aa0dc62e8b2@isocpp.org>
 <7003ddf4-5c71-4aa7-bd13-5710d8579b4c@isocpp.org> <CANh-dXmgBhfu5DE3FBrysT1VoTQ5fyQGR5VuYnmBGa8c7uFZUQ@mail.gmail.com>
 <edbb6ef4-ea8f-4f74-8492-4e52816c5ef3@isocpp.org>
 <CANh-dX=83qSJEwoE4zK0DXMESdnuFFaYVUSJJF5w1fzhF4eV7A@mail.gmail.com>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: multipart/alternative; 
	boundary="----=_Part_9_29971111.1400180560357"
X-Trace: ger.gmane.org 1400180569 18574 80.91.229.3 (15 May 2014 19:02:49 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Thu, 15 May 2014 19:02:49 +0000 (UTC)
To: std-proposals@isocpp.org
Original-X-From: std-proposals+bncBC2MLAWQ6ANBBUU62SNQKGQEXLVNNZY@isocpp.org Thu May 15 21:02:45 2014
Return-path: <std-proposals+bncBC2MLAWQ6ANBBUU62SNQKGQEXLVNNZY@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-ob0-f197.google.com ([209.85.214.197])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBC2MLAWQ6ANBBUU62SNQKGQEXLVNNZY@isocpp.org>)
	id 1Wl0vT-000537-SU
	for gclcip-std-proposals@m.gmane.org; Thu, 15 May 2014 21:02:44 +0200
Original-Received: by mail-ob0-f197.google.com with SMTP id vb8sf7158957obc.0
        for <gclcip-std-proposals@m.gmane.org>; Thu, 15 May 2014 12:02:43 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=googlemail.com; s=20120113;
        h=date:from:to:message-id:in-reply-to:references:subject:mime-version
         :x-original-sender:reply-to:precedence:mailing-list:list-id
         :list-post:list-help:list-archive:list-subscribe:list-unsubscribe
         :content-type;
        bh=HDfTQu9UCmd9CmOyNptcOQyfJuoF3rRJOeJGlmqv9Qo=;
        b=nVP+4l9XLCrBMcEafa7EElhv3NhfL3HlWyJFG6CT1WZ5sUTJ5EwGV8oWIRZF8xknkI
         Oqsf820Tsb8hde9lcXzi7sdYXBEKdtKsD5R396MNqyPWIaewctbiH/hEeU9/43MhYSlQ
         4NZyTtQetwBty93MqkoyEZ16Ju+f1DRug7ij1+Ans+DJWptsEjtmBpo94+txY+O2gJB4
         fRJsEU+g2n8wx+qGls9PwaXu3Rf9amKX+heHlyKSAqYZ6A04iyOhMuBeQ2X2I9uuZ4io
         QvZYmsCmgqEXWqFXChS9dBiy+KicHXoDGiWjFBcHDnDZrGKhsaH5sS39pvqvrP/9tgUt
         7OOA==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:date:from:to:message-id:in-reply-to:references
         :subject:mime-version:x-original-sender:reply-to:precedence
         :mailing-list:list-id:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe:content-type;
        bh=HDfTQu9UCmd9CmOyNptcOQyfJuoF3rRJOeJGlmqv9Qo=;
        b=WK+hHyC4dK1Nowg5gIXqwwNnckpRkKgvHsvqsC+akrOC11e9oxdzHj+dzh0eBAKv/5
         VXiHqqJfnyPGXNHRrimVFGQuH9WfkyWJzK/JSvS4fi81jG8g1wxSdVdbtafuSxFqTdyw
         qSH22uSNFcqy2u+d2g89xBdbUa6732J5EebJk036fl32tyBTXqn3U+4O3lHfUnBCpIj5
         u/ckyoSq0Wm0SFueyJW2vxgmugpkRy8032YimphWUtkR1doppTmx0TuLpORElvxjJeID
         069CDZDemhipNgdIMos28j0veDQ5Z7Qz1FC/Zi6kQwhwUHMQw0boWykQQA+XTPHsfn6W
         xvxw==
X-Gm-Message-State: ALoCoQkaaHTWlhgj6hKgNP5X0Rsr9HVVgTUtBgboeqlVS0k9CEkLPXobAn66GLf0Nba/NIl5x+8t
X-Received: by 10.182.104.74 with SMTP id gc10mr5891725obb.40.1400180562856;
        Thu, 15 May 2014 12:02:42 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.140.83.50 with SMTP id i47ls451243qgd.38.gmail; Thu, 15 May
 2014 12:02:42 -0700 (PDT)
X-Received: by 10.140.97.166 with SMTP id m35mr44463qge.28.1400180562187;
        Thu, 15 May 2014 12:02:42 -0700 (PDT)
In-Reply-To: <CANh-dX=83qSJEwoE4zK0DXMESdnuFFaYVUSJJF5w1fzhF4eV7A@mail.gmail.com>
X-Original-Sender: diggsey@googlemail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:10647
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/10647>

------=_Part_9_29971111.1400180560357
Content-Type: text/plain; charset=UTF-8

That may be so, but it's still better to specify the more general version 
in the standard - it's trivial for an implementation to specialize it to 
make it extra fast on UTF8, or UTF16, but it's impossible to go the other 
way. It's also less effort: regardless of how many specialisations of an 
algorithm there are, the standard just has to describe a single generic 
version which behaves exactly as the unicode standard specifies.

Even if you use only one encoding throughout your program, what if it 
happens to not be the one which all the unicode algorithms were written for?

On Thursday, 15 May 2014 19:27:17 UTC+1, Jeffrey Yasskin wrote:
>
> On Thu, May 15, 2014 at 11:23 AM, Diggory Blake <dig...@googlemail.com<javascript:>> 
> wrote: 
> > On Thursday, 15 May 2014 17:20:30 UTC+1, Jeffrey Yasskin wrote: 
> >> 
> >> On Thu, May 15, 2014 at 2:10 AM,  <pec...@gmail.com> wrote: 
> >> > I guess if neither basic_string nor char_traits can be modified your 
> >> > previous example would look like this: 
> >> > 
> >> > std::transform(lbegin_from<utf8>(input), lend_from<utf8>(input), 
> >> > lback_inserter<utf8>(result), to_upper); 
> >> 
> >> This is why we shouldn't be trying to design Unicode support in the 
> >> C++ group. Amateurs tend to think that unicode-aware algorithms can 
> >> run a codepoint at a time, while they almost always have to run a 
> >> whole string at a time. For example, 
> >> ftp://ftp.unicode.org/Public/UCD/latest/ucd/SpecialCasing.txt lists a 
> >> bunch of case conversions where the context of the character matters. 
> > 
> > 
> > That's a bad example, but the point is that whatever algorithms are 
> provided 
> > should be implemented in terms of code-points: yes, many algorithms 
> can't 
> > work on a single code-point at a time, but they should still use 
> code-points 
> > as their atomic unit. Otherwise you end up in the same situation as C 
> > runtime library with 20 different versions of every function for each 
> > specific case instead of one generic one. The algorithm for case 
> conversion 
> > can then be written once and work properly regardless of the encoding 
> used. 
> > Ultimately every string encoding is just a description of how to turn 
> bytes 
> > into code-points and back again, so code-points give a common ground 
> where 
> > all encodings are equal. 
>
> That also sounds good, but it turns out to be wrong again. The ICU 
> folks (Dick Sites in particular) have gotten significant speedups by 
> writing their algorithms using state machines directly on top of 
> encoded utf-8 and utf-16 data. You convert your data to either utf-8 
> or utf-16 when it comes into your system, and then you run algorithms 
> on the single encoding you use. It's a fool's errand to keep lots of 
> encodings inside your system. 
>

-- 

--- 
You received this message because you are subscribed to the Google Groups "ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an email to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposals/.

------=_Part_9_29971111.1400180560357
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">That may be so, but it's still better to specify the more =
general version in the standard - it's trivial for an implementation to spe=
cialize it to make it extra fast on UTF8, or UTF16, but it's impossible to =
go the other way. It's also less effort: regardless of how many specialisat=
ions of an algorithm there are, the standard just has to describe a single =
generic version which behaves exactly as the unicode standard specifies.<br=
><br>Even if you use only one encoding throughout your program, what if it =
happens to not be the one which all the unicode algorithms were written for=
?<br><br>On Thursday, 15 May 2014 19:27:17 UTC+1, Jeffrey Yasskin  wrote:<b=
lockquote class=3D"gmail_quote" style=3D"margin: 0;margin-left: 0.8ex;borde=
r-left: 1px #ccc solid;padding-left: 1ex;">On Thu, May 15, 2014 at 11:23 AM=
, Diggory Blake &lt;<a href=3D"javascript:" target=3D"_blank" gdf-obfuscate=
d-mailto=3D"1Md1iXxz6fwJ" onmousedown=3D"this.href=3D'javascript:';return t=
rue;" onclick=3D"this.href=3D'javascript:';return true;">dig...@googlemail.=
com</a>&gt; wrote:
<br>&gt; On Thursday, 15 May 2014 17:20:30 UTC+1, Jeffrey Yasskin wrote:
<br>&gt;&gt;
<br>&gt;&gt; On Thu, May 15, 2014 at 2:10 AM, &nbsp;&lt;<a>pec...@gmail.com=
</a>&gt; wrote:
<br>&gt;&gt; &gt; I guess if neither basic_string nor char_traits can be mo=
dified your
<br>&gt;&gt; &gt; previous example would look like this:
<br>&gt;&gt; &gt;
<br>&gt;&gt; &gt; std::transform(lbegin_from&lt;<wbr>utf8&gt;(input), lend_=
from&lt;utf8&gt;(input),
<br>&gt;&gt; &gt; lback_inserter&lt;utf8&gt;(result), to_upper);
<br>&gt;&gt;
<br>&gt;&gt; This is why we shouldn't be trying to design Unicode support i=
n the
<br>&gt;&gt; C++ group. Amateurs tend to think that unicode-aware algorithm=
s can
<br>&gt;&gt; run a codepoint at a time, while they almost always have to ru=
n a
<br>&gt;&gt; whole string at a time. For example,
<br>&gt;&gt; <a href=3D"ftp://ftp.unicode.org/Public/UCD/latest/ucd/Special=
Casing.txt" target=3D"_blank" onmousedown=3D"this.href=3D'ftp://ftp.unicode=
..org/Public/UCD/latest/ucd/SpecialCasing.txt';return true;" onclick=3D"this=
..href=3D'ftp://ftp.unicode.org/Public/UCD/latest/ucd/SpecialCasing.txt';ret=
urn true;">ftp://ftp.unicode.org/Public/<wbr>UCD/latest/ucd/SpecialCasing.<=
wbr>txt</a> lists a
<br>&gt;&gt; bunch of case conversions where the context of the character m=
atters.
<br>&gt;
<br>&gt;
<br>&gt; That's a bad example, but the point is that whatever algorithms ar=
e provided
<br>&gt; should be implemented in terms of code-points: yes, many algorithm=
s can't
<br>&gt; work on a single code-point at a time, but they should still use c=
ode-points
<br>&gt; as their atomic unit. Otherwise you end up in the same situation a=
s C
<br>&gt; runtime library with 20 different versions of every function for e=
ach
<br>&gt; specific case instead of one generic one. The algorithm for case c=
onversion
<br>&gt; can then be written once and work properly regardless of the encod=
ing used.
<br>&gt; Ultimately every string encoding is just a description of how to t=
urn bytes
<br>&gt; into code-points and back again, so code-points give a common grou=
nd where
<br>&gt; all encodings are equal.
<br>
<br>That also sounds good, but it turns out to be wrong again. The ICU
<br>folks (Dick Sites in particular) have gotten significant speedups by
<br>writing their algorithms using state machines directly on top of
<br>encoded utf-8 and utf-16 data. You convert your data to either utf-8
<br>or utf-16 when it comes into your system, and then you run algorithms
<br>on the single encoding you use. It's a fool's errand to keep lots of
<br>encodings inside your system.
<br></blockquote></div>

<p></p>

-- <br />
<br />
--- <br />
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br />
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to <a href=3D"mailto:std-proposals+unsubscribe@isocpp.org">std-proposa=
ls+unsubscribe@isocpp.org</a>.<br />
To post to this group, send email to <a href=3D"mailto:std-proposals@isocpp=
..org">std-proposals@isocpp.org</a>.<br />
Visit this group at <a href=3D"http://groups.google.com/a/isocpp.org/group/=
std-proposals/">http://groups.google.com/a/isocpp.org/group/std-proposals/<=
/a>.<br />

------=_Part_9_29971111.1400180560357--

.
