220 9038 <52F000AD.80502@knejp.de> article
Path: news.gmane.org!not-for-mail
From: Miro Knejp <miro@knejp.de>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Re: String to T conversions - getting it right
 this time
Date: Mon, 03 Feb 2014 21:48:45 +0100
Lines: 237
Approved: news@gmane.org
Message-ID: <52F000AD.80502@knejp.de>
References: <097fb6c8-56f9-433e-a2f0-8c0f69609bf0@isocpp.org> <lc64t2$n0m$1@ger.gmane.org> <34c26aca-59cf-4254-9ab1-31a403d4a3db@isocpp.org> <lcbben$mum$1@ger.gmane.org> <52E95CDE.40007@knejp.de> <2b93e8bc-ad3a-4c30-8f3e-30c3dd37d080@isocpp.org>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: multipart/alternative;
 boundary="------------090702090608060005020109"
X-Trace: ger.gmane.org 1391460516 10091 80.91.229.3 (3 Feb 2014 20:48:36 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Mon, 3 Feb 2014 20:48:36 +0000 (UTC)
To: std-proposals@isocpp.org
Original-X-From: std-proposals+bncBC6ONSXJ54LBBKUBYCLQKGQER4NCWEQ@isocpp.org Mon Feb 03 21:48:44 2014
Return-path: <std-proposals+bncBC6ONSXJ54LBBKUBYCLQKGQER4NCWEQ@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-la0-f71.google.com ([209.85.215.71])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBC6ONSXJ54LBBKUBYCLQKGQER4NCWEQ@isocpp.org>)
	id 1WAQRg-0003HK-1A
	for gclcip-std-proposals@m.gmane.org; Mon, 03 Feb 2014 21:48:44 +0100
Original-Received: by mail-la0-f71.google.com with SMTP id c6sf16133661lan.6
        for <gclcip-std-proposals@m.gmane.org>; Mon, 03 Feb 2014 12:48:43 -0800 (PST)
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:message-id:date:from:user-agent:mime-version:to
         :subject:references:in-reply-to:x-original-sender
         :x-original-authentication-results:reply-to:precedence:mailing-list
         :list-id:list-post:list-help:list-archive:list-subscribe
         :list-unsubscribe:content-type;
        bh=zDREOwHUlK45l59HwgpKjLebroei+1H7BF9+oiRogCg=;
        b=TE7ZoWyo5C3XrWC14S11DixMphrf5xgZWW420P9xdHUM80Q2xIgQe4VuYiA7eSh7vw
         SwhdFFBgmO/zzEPBPf/rNqZ+HizRB5pPA3NG55Y7/nK4eEwPs8siuqppIqCu5n/kuQ4Z
         yyOqoQ45+0lGwzznTUQ7vjddR/6v4lWjKY0Im9c2kkRYQOCPH15KNXJHjm3Gx2ahyiHu
         a9IUIXZ9ZcWOVbm90PTb/irLtEyI1yjdxazHwe9sq52YWOWp4l5WsGqX5DLg/EOKNuvq
         GKHOMvMH0t9vJCceNysIJ/sYyw6WRLof4juFhMjy9yHv2wJ1rZSQD/ucoHT4t7n4ubW4
         B5Zw==
X-Gm-Message-State: ALoCoQkDW1rNiusZ+aIhggBHxpRmuwSr2Sl2E8muY2qzlc9Rgnhk3fkxcgjPFvLbAMt5hwma6jhi
X-Received: by 10.152.210.197 with SMTP id mw5mr17682510lac.5.1391460523364;
        Mon, 03 Feb 2014 12:48:43 -0800 (PST)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.180.11.41 with SMTP id n9ls496851wib.5.gmail; Mon, 03 Feb 2014
 12:48:42 -0800 (PST)
X-Received: by 10.14.209.129 with SMTP id s1mr46223297eeo.21.1391460522107;
        Mon, 03 Feb 2014 12:48:42 -0800 (PST)
Original-Received: from mail-out.m-online.net (mail-out.m-online.net. [212.18.0.9])
        by mx.google.com with ESMTPS id 43si37686153eeh.220.2014.02.03.12.48.41
        for <std-proposals@isocpp.org>
        (version=TLSv1 cipher=RC4-SHA bits=128/128);
        Mon, 03 Feb 2014 12:48:42 -0800 (PST)
Received-SPF: neutral (google.com: 212.18.0.9 is neither permitted nor denied by best guess record for domain of miro@knejp.de) client-ip=212.18.0.9;
Original-Received: from frontend1.mail.m-online.net (unknown [192.168.8.180])
	by mail-out.m-online.net (Postfix) with ESMTP id 3fJ1N95r5Kz4KK4C
	for <std-proposals@isocpp.org>; Mon,  3 Feb 2014 21:48:41 +0100 (CET)
Original-Received: from www.knejp.de (ppp-188-174-143-74.dynamic.mnet-online.de [188.174.143.74])
	(using TLSv1 with cipher DHE-RSA-AES256-SHA (256/256 bits))
	(No client certificate requested)
	by mail.mnet-online.de (Postfix) with SMTP id 3fJ1N92vckzbbcY
	for <std-proposals@isocpp.org>; Mon,  3 Feb 2014 21:48:41 +0100 (CET)
Original-Received: from [192.168.42.4] ([192.168.42.4])
	by www.knejp.de
	; Mon, 3 Feb 2014 21:48:37 +0100
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.2.0
In-Reply-To: <2b93e8bc-ad3a-4c30-8f3e-30c3dd37d080@isocpp.org>
X-Original-Sender: miro@knejp.de
X-Original-Authentication-Results: mx.google.com;       spf=neutral
 (google.com: 212.18.0.9 is neither permitted nor denied by best guess record
 for domain of miro@knejp.de) smtp.mail=miro@knejp.de
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:9038
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/9038>

This is a multi-part message in MIME format.
--------------090702090608060005020109
Content-Type: text/plain; charset=UTF-8; format=flowed


Am 03.02.2014 12:18, schrieb Olaf van der Spek:
> On Wednesday, January 29, 2014 8:56:14 PM UTC+1, Miro Knejp wrote:
>
>     I am now using the following interface in the format parser
>     implementation:
>
>     pair<optional<T>, Iter> parse_integer<T>(Iter first, Iter last, int
>     radix = 10)
>
>     and the convenience overload
>
>     optional<T> parse_integer<T>(string_view s, int radix = 10);
>
>     which could, using internal tag dispatching, be reduced to
>
>     parse<T>(...)
>
>     The signatures are very easy to use and give me all I need. Both 
>
>
>  Do you perhaps have a link to a project using this interface?
>
The format implementation can be found here: 
https://github.com/mknejp/std-format
The defintions of the parse_xxx methods (of which there currently are 
only integer versions) are in include/std-format/detail/parse_tools.hpp 
and they are used in include/std-format/detail/format_parser.hpp.

Not sure how serious an example that is as for processing the actual 
format string I only need to parse integers at two locations in 
format_parser.hpp and there is no error reporting (only 
success/failure). However, at some point I also need to start parsing 
the format options for all the builtin and std types and there it will 
become clearer how useful the interface really is. Last time I had to 
write number parsing myself was some 10 years ago so please don't mind 
if the implementation of the parse methods isn't perfect.

Considering the debate about octal numbers and prefixes I went along and 
split the methods up depending on use case, so I have:

parse_integer(...) <- accepts [+-]?[0-9a-zA-Z]+

parse_radix_prefix(...) <- accepts (0[xXbX]?)? retuning the radix 2, 8, 
16 or 0 if the pattern doesn't apply

parse_prefixed_integer(...) <- accepts [+-]?(0[xXbB]?)?[0-9]+ and if the 
prefix is not recognized uses the radix passed as argument

The actual range of valid characters in [0-9a-zA-Z] depends on the radix 
and the minus sign is only accepted for signed integer types. None of 
them skip any whitespace characters on any end of the string. They 
consume all valid characters even if overflow occurs. Feel free to 
replace any character with culture specific signs and digits when 
locales apply.

Then I was thinking some more about the return/error dilemma. Inspired 
by the mentioning of match_integer three use case scenarios come to mind:

 1. Parsing a longer text. At some point you determine that at position
    i should be a number. This is the case where you probably need the
    most information: success/failure, error description if it failed
    and in both cases an iterator to the next position so you can
    continue processing the remaining source. I guess this is what I
    tried to cover with my parse_xxx interface.
 2. You have a string and it hast to contain a number. In this case the
    source has to match exactly with zero tolerance. This would be the
    case for a match_xxx interface where invalid characters at the end
    cause a failed conversion.
 3. You have a string of some description and expect it to begin with a
    number. Invalid characters at the end of the source range do not
    cause an error. This might involve skipping of whitespaces before
    the number. I see this use case occurring especially in exercises
    and introductory courses when dealing with basic user input for the
    first time. I don't really have a fitting name for such an
    interface. from_string maybe? No idea.

I see number 1 as the most fundamental. The interface is based on 
iterators and can thus work with almost any type of input. 2 and 3 can 
be implemented in terms of 1 and a string_view overload would probably 
be used more often there. As soon as locales are involved all three 
should automatically recognize the correct grouping, separating and 
decimal characters.

-- 

--- 
You received this message because you are subscribed to the Google Groups "ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an email to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposals/.

--------------090702090608060005020109
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<html>
  <head>
    <meta content=3D"text/html; charset=3DUTF-8" http-equiv=3D"Content-Type=
">
  </head>
  <body text=3D"#000000" bgcolor=3D"#FFFFFF">
    <br>
    <div class=3D"moz-cite-prefix">Am 03.02.2014 12:18, schrieb Olaf van
      der Spek:<br>
    </div>
    <blockquote
      cite=3D"mid:2b93e8bc-ad3a-4c30-8f3e-30c3dd37d080@isocpp.org"
      type=3D"cite">
      <div dir=3D"ltr">On Wednesday, January 29, 2014 8:56:14 PM UTC+1,
        Miro Knejp wrote:
        <blockquote class=3D"gmail_quote" style=3D"margin: 0;margin-left:
          0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;">I am now
          using the following interface in the format parser
          implementation:
          <br>
          <br>
          pair&lt;optional&lt;T&gt;, Iter&gt;
          parse_integer&lt;T&gt;(Iter first, Iter last, int <br>
          radix =3D 10)
          <br>
          <br>
          and the convenience overload
          <br>
          <br>
          optional&lt;T&gt; parse_integer&lt;T&gt;(string_view s, int
          radix =3D 10);
          <br>
          <br>
          which could, using internal tag dispatching, be reduced to
          <br>
          <br>
          parse&lt;T&gt;(...)
          <br>
          <br>
          The signatures are very easy to use and give me all I need.
          Both </blockquote>
        <div><br>
        </div>
        <div>=C2=A0Do you perhaps have a link to a project using this
          interface?</div>
        <div><br>
        </div>
      </div>
    </blockquote>
    The format implementation can be found here:
    <a class=3D"moz-txt-link-freetext" href=3D"https://github.com/mknejp/st=
d-format">https://github.com/mknejp/std-format</a><br>
    The defintions of the parse_xxx methods (of which there currently
    are only integer versions) are in
    include/std-format/detail/parse_tools.hpp and they are used in
    include/std-format/detail/format_parser.hpp.<br>
    <br>
    Not sure how serious an example that is as for processing the actual
    format string I only need to parse integers at two locations in
    format_parser.hpp and there is no error reporting (only
    success/failure). However, at some point I also need to start
    parsing the format options for all the builtin and std types and
    there it will become clearer how useful the interface really is.
    Last time I had to write number parsing myself was some 10 years ago
    so please don't mind if the implementation of the parse methods
    isn't perfect.<br>
    <br>
    Considering the debate about octal numbers and prefixes I went along
    and split the methods up depending on use case, so I have:<br>
    <br>
    parse_integer(...) &lt;- accepts [+-]?[0-9a-zA-Z]+<br>
    <br>
    parse_radix_prefix(...) &lt;- accepts (0[xXbX]?)? retuning the radix
    2, 8, 16 or 0 if the pattern doesn't apply<br>
    <br>
    parse_prefixed_integer(...) &lt;- accepts [+-]?(0[xXbB]?)?[0-9]+ and
    if the prefix is not recognized uses the radix passed as argument<br>
    <br>
    The actual range of valid characters in [0-9a-zA-Z] depends on the
    radix and the minus sign is only accepted for signed integer types.
    None of them skip any whitespace characters on any end of the
    string. They consume all valid characters even if overflow occurs.
    Feel free to replace any character with culture specific signs and
    digits when locales apply.<br>
    <br>
    Then I was thinking some more about the return/error dilemma.
    Inspired by the mentioning of match_integer three use case scenarios
    come to mind:<br>
    <ol>
      <li>Parsing a longer text. At some point you determine that at
        position i should be a number. This is the case where you
        probably need the most information: success/failure, error
        description if it failed and in both cases an iterator to the
        next position so you can continue processing the remaining
        source. I guess this is what I tried to cover with my parse_xxx
        interface.<br>
      </li>
      <li>You have a string and it hast to contain a number. In this
        case the source has to match exactly with zero tolerance. This
        would be the case for a match_xxx interface where invalid
        characters at the end cause a failed conversion.</li>
      <li>You have a string of some description and expect it to begin
        with a number. Invalid characters at the end of the source range
        do not cause an error. This might involve skipping of
        whitespaces before the number. I see this use case occurring
        especially in exercises and introductory courses when dealing
        with basic user input for the first time. I don't really have a
        fitting name for such an interface. from_string maybe? No idea.<br>
      </li>
    </ol>
    <p>I see number 1 as the most fundamental. The interface is based on
      iterators and can thus work with almost any type of input. 2 and 3
      can be implemented in terms of 1 and a string_view overload would
      probably be used more often there. As soon as locales are involved
      all three should automatically recognize the correct grouping,
      separating and decimal characters.<br>
    </p>
  </body>
</html>

<p></p>

-- <br />
&nbsp;<br />
--- <br />
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br />
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to std-proposals+unsubscribe@isocpp.org.<br />
To post to this group, send email to std-proposals@isocpp.org.<br />
Visit this group at <a href=3D"http://groups.google.com/a/isocpp.org/group/=
std-proposals/">http://groups.google.com/a/isocpp.org/group/std-proposals/<=
/a>.<br />

--------------090702090608060005020109--


.
