220 26427 <fa7ab47a-a3c7-4d23-a5a1-e13654b0c8c6@isocpp.org> article
Path: news.gmane.org!not-for-mail
From: Jean-Marc Bourguet <jm.bourguet@gmail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: iswalpha and locales
Date: Mon, 27 Jun 2016 08:18:42 -0700 (PDT)
Lines: 133
Approved: news@gmane.org
Message-ID: <fa7ab47a-a3c7-4d23-a5a1-e13654b0c8c6@isocpp.org>
References: <5057e854-b2ea-4c39-81c7-367bc3e54080@isocpp.org>
 <5458114.SXJFP9uSZR@tjmaciei-mobl1>
 <op.yjmk55nhhnjspo@debian>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: multipart/mixed; 
	boundary="----=_Part_628_444703992.1467040722386"
X-Trace: ger.gmane.org 1467040727 24961 80.91.229.3 (27 Jun 2016 15:18:47 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Mon, 27 Jun 2016 15:18:47 +0000 (UTC)
Cc: asorenji@gmail.com
To: ISO C++ Standard - Future Proposals <std-proposals@isocpp.org>
Original-X-From: std-proposals+bncBC56XVMTUYPRBU4HYW5QKGQETIDJZBQ@isocpp.org Mon Jun 27 17:18:47 2016
Return-path: <std-proposals+bncBC56XVMTUYPRBU4HYW5QKGQETIDJZBQ@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-pf0-f197.google.com ([209.85.192.197])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBC56XVMTUYPRBU4HYW5QKGQETIDJZBQ@isocpp.org>)
	id 1bHYJC-000661-2q
	for gclcip-std-proposals@m.gmane.org; Mon, 27 Jun 2016 17:18:46 +0200
Original-Received: by mail-pf0-f197.google.com with SMTP id e189sf405745545pfa.2
        for <gclcip-std-proposals@m.gmane.org>; Mon, 27 Jun 2016 08:18:45 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=isocpp-org.20150623.gappssmtp.com; s=20150623;
        h=date:from:to:cc:message-id:in-reply-to:references:subject
         :mime-version:x-original-sender:reply-to:precedence:mailing-list
         :list-id:x-spam-checked-in-group:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe;
        bh=16w0aUde1/sD4FYeM8F+7hs+T/FdhdtCpZHDIiMchI8=;
        b=KR3lSSvzVvoTbRyFIiP+JEbXzYhpPoTeS7qsXpYNJ7091O6tTlFC05CW31N1BVA85g
         lQARkS04AzOt7Jx830WITGi7HSP7BS72q6OaqT2dukSxpKN8DMESflFzZgRrOTxc+IvE
         Q93JUHyClh2R/uCjdvr0kSAccC6uvI1liIfh+kf/rz5spVAiCyV5bJ7UBLZgtSPdgvcK
         TDW6Gpja0d9FfOxCN0zbWECclaVp1HT9LxS0xsSygnL6eNvi37F1nQN2zrl3E1hF19CT
         nIYoyoBJGxu6An9+86NEmRsN2n2GEtfeJm6xhyBAehzVmOUEsDVbpJhuiOdN9okwXCTc
         zdpA==
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=gmail.com; s=20120113;
        h=date:from:to:cc:message-id:in-reply-to:references:subject
         :mime-version:x-original-sender:reply-to:precedence:mailing-list
         :list-id:x-spam-checked-in-group:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe;
        bh=16w0aUde1/sD4FYeM8F+7hs+T/FdhdtCpZHDIiMchI8=;
        b=iU9iqHLYEcG3rySVbgPBTQ9UR3a1cJaOD1YM6MN9N+d/AHjO4FMqGLoPBZYJZYAZ/n
         UI0DIy2n/GsM52EyMxE7djqjEsd5Fke8Tojud2wNYDi/n4mT1/MTZmuahpNWtBfp1K1y
         ZtsncoEhtYR2cxYqhMzg1fpk9nuVjfEB2D7mmVgads8JdGaEoxAh8eg4OvP5z5R2O1pI
         NC3YyEnW19SEGk6XkBO081085dKccimwh29mxQlnY08jK+Zm5Qi3E7w874zO1fyoXlNI
         zpyHrxvW/gNkhtohBZylkVkOTBrBVXMGjUOvoFUry4aJUs5ncBTsGpDkIGS2dbLde4BF
         vzvQ==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:date:from:to:cc:message-id:in-reply-to
         :references:subject:mime-version:x-original-sender:reply-to
         :precedence:mailing-list:list-id:x-spam-checked-in-group:list-post
         :list-help:list-archive:list-subscribe:list-unsubscribe;
        bh=16w0aUde1/sD4FYeM8F+7hs+T/FdhdtCpZHDIiMchI8=;
        b=MlWgRbwPwNeyEfHUaT+dqo7XnAVrjJ1BKrBXlrxlWMpx1kb+f3KzX6f5aXmywEoGI6
         uUoa10HnQwelo1JL788gaFDvH+y6fK/5NiOYzer1bVEP4NC0nA2VlepxCdUHBYYqqwx4
         B8cuNyivb9VB2eHsE5ghwDket7KtsOFiicONrsF93YvWrvXBjGw1KcFUHVCW2EjjwY2f
         zQVlvY8+hrjjx4aHgORh9wMFm54v59WHmojzJYaedFVrfUvLgbLXoJJwftGiw4VqRmK+
         YTXtv4OyCXS1hd3xOXvQvKmZDzJm/YnWOZ7xC3jrvbsZy5pkZFqhhqOqijannBD+pYH9
         M0VQ==
X-Gm-Message-State: ALyK8tJlC/88uYvWRvSVUUMpD0BpqjGZiCXZ8E7BWKut6H4q/6oGJsbPS1B6nfDpIgUeiw==
X-Received: by 10.66.138.81 with SMTP id qo17mr17950952pab.26.1467040724751;
        Mon, 27 Jun 2016 08:18:44 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.107.51.129 with SMTP id z123ls393966ioz.10.gmail; Mon, 27 Jun
 2016 08:18:43 -0700 (PDT)
X-Received: by 10.36.127.66 with SMTP id r63mr238448itc.2.1467040723432;
        Mon, 27 Jun 2016 08:18:43 -0700 (PDT)
In-Reply-To: <op.yjmk55nhhnjspo@debian>
X-Original-Sender: JM.Bourguet@gmail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Spam-Checked-In-Group: std-proposals@isocpp.org
X-Google-Group-Id: 399137483710
List-Post: <https://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <https://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <https://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <https://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>,
 <https://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:26427
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/26427>

------=_Part_628_444703992.1467040722386
Content-Type: multipart/alternative; 
	boundary="----=_Part_629_307912543.1467040722394"

------=_Part_629_307912543.1467040722394
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

Le samedi 25 juin 2016 20:19:06 UTC+2, Aso Renji a =C3=A9crit :
>
> Or, maybe wint_t is not unicode? In this case you know any other characte=
r=20
>  =20
> encoding with wide (not multi-byte) characters?=20
>
>
The encoding used for wchar_t is locale specific.  The encoding model of C=
=20
and C++ is that a locale has one charset and three encodings for that=20
charset (a narrow encoding using char as encoding unit, which can have=20
shift state and be multi-byte, a wide encoding using wchar_t, which may not=
=20
have shift state and may not use several wchar_t to represent a code-point,=
=20
an external encoding which has less restrictions -- the best known is that=
=20
end of line may be represented by something else than a single character --=
=20
which is observable by looking at difference between binary and text file=
=20
IO; if my understanding is correct, you can have a locale using UTF-8 as=20
narrow encoding, UTF-32 as wide encoding and UTF-16 with BOM at start and=
=20
CR-LF as line separator as external encoding).  I'm pretty sure -- I don't=
=20
have access to that hardware/software combination anymore -- that I've used=
=20
systems which had available at the same time:

- ascii char set, char and wchar_t are using directly the code point value

- ISO-8859-X charsets, char and wchar_t were using directly the code point=
=20
value (note that this is different from the Linux behavior which is to use=
=20
the code point value for char and the Unicode code point value for wchar_t=
=20
-- note also that this is the case I wanted to check if I was remembering=
=20
correctly)

- CJK charsets with char being a serialized form of EUC and wchar_t=20
grouping the number of chars needed for the character (EUC is a way to=20
encode a subset of ISO 2022 streams without using shift state with a=20
maximum of 4 bytes per char)

- the unicode charset, char is UTF-8 and wchar_t is the code point.

Most language/region pair had variant locales for several charset (for=20
instance Latin-1, Latin-9 and Unicode for the French locales, each of which=
=20
would have used a different wide encoding).

Yours

--=20
You received this message because you are subscribed to the Google Groups "=
ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
To view this discussion on the web visit https://groups.google.com/a/isocpp=
..org/d/msgid/std-proposals/fa7ab47a-a3c7-4d23-a5a1-e13654b0c8c6%40isocpp.or=
g.

------=_Part_629_307912543.1467040722394
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">Le samedi 25 juin 2016 20:19:06 UTC+2, Aso Renji a =C3=A9c=
rit=C2=A0:<blockquote class=3D"gmail_quote" style=3D"margin: 0;margin-left:=
 0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;">Or, maybe wint_t is =
not unicode? In this case you know any other character =C2=A0
<br>encoding with wide (not multi-byte) characters?
<br><br></blockquote><div><br></div><div>The encoding used for wchar_t is l=
ocale specific. =C2=A0The encoding model of C and C++ is that a locale has =
one charset and three encodings for that charset (a narrow encoding using c=
har as encoding unit, which can have shift state and be multi-byte, a wide =
encoding using wchar_t, which may not have shift state and may not use seve=
ral wchar_t to represent a code-point, an external encoding which has less =
restrictions -- the best known is that end of line may be represented by so=
mething else than a single character -- which is observable by looking at d=
ifference between binary and text file IO; if my understanding is correct, =
you can have a locale using UTF-8 as narrow encoding, UTF-32 as wide encodi=
ng and UTF-16 with BOM at start and CR-LF as line separator as external enc=
oding). =C2=A0I&#39;m pretty sure -- I don&#39;t have access to that hardwa=
re/software combination anymore -- that I&#39;ve used systems which had ava=
ilable at the same time:</div><div><br></div><div>- ascii char set, char an=
d wchar_t are using directly the code point value</div><div><br></div><div>=
- ISO-8859-X charsets, char and wchar_t were using directly the code point =
value (note that this is different from the Linux behavior which is to use =
the code point value for char and the Unicode code point value for wchar_t =
-- note also that this is the case I wanted to check if I was remembering c=
orrectly)</div><div><br></div><div>- CJK charsets with char being a seriali=
zed form of EUC and wchar_t grouping the number of chars needed for the cha=
racter (EUC is a way to encode a subset of ISO 2022 streams without using s=
hift state with a maximum of 4 bytes per char)</div><div><br></div><div>- t=
he unicode charset, char is UTF-8 and wchar_t is the code point.</div><div>=
<br></div><div>Most language/region pair had variant locales for several ch=
arset (for instance Latin-1, Latin-9 and Unicode for the French locales, ea=
ch of which would have used a different wide encoding).</div><div><br></div=
><div>Yours</div></div>

<p></p>

-- <br />
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br />
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to <a href=3D"mailto:std-proposals+unsubscribe@isocpp.org">std-proposa=
ls+unsubscribe@isocpp.org</a>.<br />
To post to this group, send email to <a href=3D"mailto:std-proposals@isocpp=
..org">std-proposals@isocpp.org</a>.<br />
To view this discussion on the web visit <a href=3D"https://groups.google.c=
om/a/isocpp.org/d/msgid/std-proposals/fa7ab47a-a3c7-4d23-a5a1-e13654b0c8c6%=
40isocpp.org?utm_medium=3Demail&utm_source=3Dfooter">https://groups.google.=
com/a/isocpp.org/d/msgid/std-proposals/fa7ab47a-a3c7-4d23-a5a1-e13654b0c8c6=
%40isocpp.org</a>.<br />

------=_Part_629_307912543.1467040722394--

------=_Part_628_444703992.1467040722386--

.
