220 18478 <8118aad4-4ede-43a5-9f2c-b2b483cf0e02@isocpp.org> article
Path: news.gmane.org!not-for-mail
From: Nicol Bolas <jmckesson@gmail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Exposing string small string optimization size to
 avoid heap allocation.
Date: Sat, 6 Jun 2015 23:57:25 -0700 (PDT)
Lines: 299
Approved: news@gmane.org
Message-ID: <8118aad4-4ede-43a5-9f2c-b2b483cf0e02@isocpp.org>
References: <19736033-58d4-441d-ae73-19804f8ef575@isocpp.org>
 <a560b06f-1e11-4c88-ba5e-db3101a4c86f@isocpp.org>
 <5010db84-4ed4-4d52-a9da-10b8bde377c6@isocpp.org>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: multipart/mixed; 
	boundary="----=_Part_1119_157879467.1433660245863"
X-Trace: ger.gmane.org 1433660263 13762 80.91.229.3 (7 Jun 2015 06:57:43 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Sun, 7 Jun 2015 06:57:43 +0000 (UTC)
Cc: german.diago@hubblehome.com
To: std-proposals@isocpp.org
Original-X-From: std-proposals+bncBCEKFTV6ZUMBBVWWZ6VQKGQEFPYLK3I@isocpp.org Sun Jun 07 08:57:29 2015
Return-path: <std-proposals+bncBCEKFTV6ZUMBBVWWZ6VQKGQEFPYLK3I@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-pa0-f71.google.com ([209.85.220.71])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBCEKFTV6ZUMBBVWWZ6VQKGQEFPYLK3I@isocpp.org>)
	id 1Z1UWP-0001i6-0t
	for gclcip-std-proposals@m.gmane.org; Sun, 07 Jun 2015 08:57:29 +0200
Original-Received: by pabqy3 with SMTP id qy3sf197321049pab.3
        for <gclcip-std-proposals@m.gmane.org>; Sat, 06 Jun 2015 23:57:27 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=gmail.com; s=20120113;
        h=date:from:to:cc:message-id:in-reply-to:references:subject
         :mime-version:content-type:x-original-sender:reply-to:precedence
         :mailing-list:list-id:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe;
        bh=3nVSTrJjTB15AXbYntb0wCvrNvyhq+NQyvSEMzya5oU=;
        b=SvYwCsz1FoelrhnAbLH7tXm0GkHDqA/AZHm4MQzA2zjc6Q+tj5n28QRCtcYZ+Dz9oP
         2vn0FzUVwUMxYH5dR0bX7GOVkiJvW4IjVHuQ3odBQ4He+z5hoigztnhBQcXrJ+okU2K6
         91rKP/jWfATE2KbOeq/5ZTp8Be4e2K6VkiB0i5JEnU6mpj3gRJlnfk1px2y6Pk6OqptZ
         Cd+apCdg7OVu5OrBdIhB6BAb9HuJdQrKTe3mjm825WTHjUo2cjESIifU6g12azKdCP5H
         bpj5uYS9BkSGcgPMT3QcJ9//XXUruCnwZQ0FFfCl5Y1ryWIOEVthK+A+DH8ZV0E2oiAI
         I6Pw==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:date:from:to:cc:message-id:in-reply-to
         :references:subject:mime-version:content-type:x-original-sender
         :reply-to:precedence:mailing-list:list-id:list-post:list-help
         :list-archive:list-subscribe:list-unsubscribe;
        bh=3nVSTrJjTB15AXbYntb0wCvrNvyhq+NQyvSEMzya5oU=;
        b=k2tTctmNm4oigbsZdwrPc88hPs47FSrr/L1OlFV83cauvKWfY35VUqlcE7ofMCZx5a
         GusDKmDwG86mihAhq+llU4XtOFoylqnGrAaJzUJC+Qw3rkPbuqB1d5n1qMLZUIw+TZcQ
         DTwewDntP+cGDOZPQWtEZoaSLAjdPRWI9u2eGqk3Q7j/wVY2RSMG9jBw3qw+RRK72FnO
         Qloe4kHTraRpF9tMe6f0kz5JMhkjWmv/FHDWnh/WlyjjFVJDYMe99ZZ7T5iUMREoRH6J
         PnhpY7nlaBUhTnSqocJpKtmxQOc72uCl5LHvXRii6tLfgZPxUQi/H7aMY+8A6NDIrBXR
         1TTw==
X-Gm-Message-State: ALoCoQmgMLkR7XbNhWKR4N3ETiYvhPg701Z8vk6AKvZxxGqzKiY6X9zU9v1o91bZ1fHou/tyBVc1
X-Received: by 10.66.149.35 with SMTP id tx3mr15160870pab.24.1433660247754;
        Sat, 06 Jun 2015 23:57:27 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.140.82.38 with SMTP id g35ls2135209qgd.7.gmail; Sat, 06 Jun
 2015 23:57:26 -0700 (PDT)
X-Received: by 10.140.28.73 with SMTP id 67mr134746qgy.36.1433660246719;
        Sat, 06 Jun 2015 23:57:26 -0700 (PDT)
In-Reply-To: <5010db84-4ed4-4d52-a9da-10b8bde377c6@isocpp.org>
X-Original-Sender: jmckesson@gmail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>,
 <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:18478
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/18478>

------=_Part_1119_157879467.1433660245863
Content-Type: multipart/alternative; 
	boundary="----=_Part_1120_1337897824.1433660245863"

------=_Part_1120_1337897824.1433660245863
Content-Type: text/plain; charset=UTF-8

On Sunday, June 7, 2015 at 2:08:20 AM UTC-4, Sean Middleditch wrote:
>
> The real issue with SSO, in my experience, is a bit more complicated.
>
> What if the std::string implementation doesn't support SSO at the string 
> lengths you need? Detecting it would at best just let you add a 
> static_assert telling users that the implementation is unusable. It doesn't 
> actually fix the problem with the implementation; it just makes it easier 
> to identify. And once you've rewritten std::string for one implementation, 
> you might as well just use your version on every implementation so you get 
> reliable behavior no matter where you go.
>
> If the standard didn't rely on "QoI" here and just set something users can 
> rely on, we'd all be in a much better place. Just like how graphics APIs 
> provide implementation-provided upper bounds for many values but mandate a 
> reliable cross-implementation lower-bound, which is essential to writing 
> graphics programs that actually work across multiple 
> implementations/drivers.
>

OK, imagine if, in 1998, the standards committee took your advice. Let's 
say they decided to formalize a "QoI" feature of std::basic_string by, not 
merely allowing it, but *requiring* it.

Only that optimization was copy-on-write, not SSO.

Remember: back then, we didn't know SSO was a good idea. Back then, CoW 
sounded far more like the reasonable solution to so many string problems. 
It was quite an optimization.

Back then.

The world has changed a lot in the 20+ years since then. What happens if 
another shift takes place, and SSO goes from being an optimization to a 
pessimization like CoW did? Now, all your standardization of that 
implementation does is force the string type to be weakened.

The standard's committee should not be in the business of predicting the 
future. Especially considering how bad they are at it.
 

> An even more general solution would allow the user to specify a traits 
> type that defined that minimal bound. When we write our own STL 
> replacements we typically don't bother with such traits, since we can just 
> design to our specific needs, but the standard library probably wants to be 
> a bit more widely applicable.
>

The standard library is already "widely applicable". That it is not 
*universally* applicable is not a bug; it's a feature. It allows the 
standard library to be relatively simple to use, to not have a bunch of 
arbitrary tweaks for minor things that 99% of users won't care about.

The last thing we need is a bunch of arbitrary policies added on to our 
types.

On Sunday, June 7, 2015 at 2:21:50 AM UTC-4, Sean Middleditch wrote:
>
> (sorry for the self-reply)
>
> An additional important point is interoperability. Just writing your own 
> string class is very problematic, even when necessary. If you need to work 
> around std::string for your environment then you also very likely need to 
> avoid libraries that use std::string. Which is most of them; that's the 
> whole point of having a standard string type. string_view will lessen the 
> problems here to a degree, but certainly not eliminate them.
>

"to a degree"? How about "mostly"?

Oh sure, `string_view` won't magically make all existing uses of 
std::string in APIs transform into std::string_view. But the option would 
now exist. And most APIs don't actually care about the *storage* for a 
string; it only needs to read the data. So most APIs that take std::string 
or some similar string can very easily have overloads that take 
std::string_view.

So, with appropriate code updates, you will have eliminated about 90% of 
such interop problems. The remaining 10% can be divided into these 
categories:

1) Functions that return std::string.

2) Functions that "return" std::string by parameter.

3) Functions that modify the contents of a std::string.

#1 and #2 can't be changed; you have to pick *some* type to use to handle 
the storage of a string. And many users will pick std::string because... 
it's there. Even if they picked the string type you like, that may not be 
the string type other people like. And trying to design an "omni_string", 
which is all things to all people, is a trail that ends in tears.

#3 generally can't be changed either. Most of such functions require the 
insertion or deletions of characters. Even seemingly simple modifications 
like upper/lowercase changes are impossible to do in Unicode without being 
able to insert/delete characters. Even in UTF-32, there are codepoint 
sequences which are considered one case, that map to a single codepoint in 
another case.

I'd say the 90% interoperability that std::string_view gives you is the 
best you can reasonably expect to get. Take it and move on.

It should also be noted that your attempt at a bunch of policies for 
std::basic_string wouldn't work for interop anyway. At least, not without 
copying. If the library uses a 16-character SSO std::string, but you use a 
32-character one, you have to copy when crossing that boundary. That's just 
as true as if it were using a different allocator on its std::string.
 

> This is something I've personally run into a lot with games, where some 
> otherwise very handy library ends up depending on a verboten STL type, or 
> Boost, or one of the oft verboten C++ features, or ends up working around 
> the same problems we have by implementing their own non-standard framework 
> library (and then we'd have to get all these libraries interoperating with 
> us and each other), and so we can't use the library and have to reimplement 
> it from scratch as well as the standard types.
>
> It's a domino effect of wasted time and money.
>

This is *your* problem, not C++'s.

Your problem is that you aren't writing C++ code; you're writing to some 
!C++ language, which is a functional subset of C++. That's all well and 
good, but the rest of us *don't* live in your little box. And if you expect 
us to do so, you'll be waiting a while.

I understand that you didn't necessarily choose your programming 
environment. But at the end of the day, you can't complain that some C++ 
library's writer dared to actually *use C++* in their code, rather than the 
arbitrary subset you're limited to.

If you are allergic to Boost, that's something I can understand. But by 
rejecting it, you are rejecting every library that uses it as a dependency. 
And that's a lot of code. That was the choice you made when you rejected 
it, and you should have considered that when making that choice (assuming 
the choice wasn't made for you). The same goes for any other "oft verboten 
C++ features" (FYI: Just because something is 'oft' for you doesn't make it 
'oft' for everyone else too).

People who actually use the language in their code is not a problem that 
the language needs to solve. It's only "a domino effect of wasted time and 
money" for those who have arbitrary limits placed on what parts of the 
language they can use. And while I empathize with you, complaining about it 
or requesting features to "correct" this "problem" is the wrong way to go 
about it.

-- 

--- 
You received this message because you are subscribed to the Google Groups "ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an email to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposals/.

------=_Part_1120_1337897824.1433660245863
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">On Sunday, June 7, 2015 at 2:08:20 AM UTC-4, Sean Middledi=
tch wrote:<blockquote class=3D"gmail_quote" style=3D"margin: 0;margin-left:=
 0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;"><div dir=3D"ltr">The=
 real issue with SSO, in my experience, is a bit more complicated.<div><br>=
</div><div>What
 if the std::string implementation doesn't support SSO at the string=20
lengths you need? Detecting it would at best just let you add a=20
static_assert telling users that the implementation is unusable. It=20
doesn't actually fix the problem with the implementation; it just makes=20
it easier to identify. And once you've rewritten std::string for one=20
implementation, you might as well just use your version on every=20
implementation so you get reliable behavior no matter where you go.</div><d=
iv><br></div><div>If
 the standard didn't rely on "QoI" here and just set something users can
 rely on, we'd all be in a much better place. Just like how graphics=20
APIs provide implementation-provided upper bounds for many values but=20
mandate a reliable cross-implementation lower-bound, which is essential=20
to writing graphics programs that actually work across multiple=20
implementations/drivers.</div></div></blockquote><div><br>OK, imagine=20
if, in 1998, the standards committee took your advice. Let's say they=20
decided to formalize a "QoI" feature of std::basic_string by, not merely al=
lowing it, but <i>requiring</i> it.<br><br>Only that optimization was copy-=
on-write, not SSO.<br><br>Remember:
 back then, we didn't know SSO was a good idea. Back then, CoW sounded=20
far more like the reasonable solution to so many string problems. It was
 quite an optimization.<br><br>Back then.<br><br>The world has changed a lo=
t in the 20+=20
years since then. What happens if another shift takes place, and SSO=20
goes from being an optimization to a pessimization like CoW did? Now,=20
all your standardization of that implementation does is force the string
 type to be weakened.<br><br>The standard's committee should not be in the =
business of predicting the future. Especially considering how bad they are =
at it.<br>&nbsp;</div><blockquote class=3D"gmail_quote" style=3D"margin: 0;=
margin-left: 0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;"><div dir=
=3D"ltr"><div>An
 even more general solution would allow the user to specify a traits=20
type that defined that minimal bound. When we write our own STL=20
replacements we typically don't bother with such traits, since we can=20
just design to our specific needs, but the standard library probably=20
wants to be a bit more widely applicable.</div></div></blockquote><br>The s=
tandard library is already "widely applicable". That it is not <i>universal=
ly</i>
 applicable is not a bug; it's a feature. It allows the standard library
 to be relatively simple to use, to not have a bunch of arbitrary tweaks
 for minor things that 99% of users won't care about.<br><br>The last thing=
 we need is a bunch of arbitrary policies added on to our types.<br><br>On =
Sunday, June 7, 2015 at 2:21:50 AM UTC-4, Sean Middleditch wrote:<blockquot=
e class=3D"gmail_quote" style=3D"margin: 0;margin-left: 0.8ex;border-left: =
1px #ccc solid;padding-left: 1ex;"><div dir=3D"ltr">(sorry for the self-rep=
ly)<div><br></div><div>An additional important point is interoperability. J=
ust writing your own string class is very problematic, even when necessary.=
 If you need to work around std::string for your environment then you also =
very likely need to avoid libraries that use std::string. Which is most of =
them; that's the whole point of having a standard string type. string_view =
will lessen the problems here to a degree, but certainly not eliminate them=
..</div></div></blockquote><div><br>"to a degree"? How about "mostly"?<br><b=
r>Oh sure, `string_view` won't magically make all existing uses of std::str=
ing in APIs transform into std::string_view. But the option would now exist=
.. And most APIs don't actually care about the <i>storage</i> for a string; =
it only needs to read the data. So most APIs that take std::string or some =
similar string can very easily have overloads that take std::string_view.<b=
r><br>So, with appropriate code updates, you will have eliminated about 90%=
 of such interop problems. The remaining 10% can be divided into these cate=
gories:<br><br>1) Functions that return std::string.<br><br>2) Functions th=
at "return" std::string by parameter.<br><br>3) Functions that modify the c=
ontents of a std::string.<br><br>#1 and #2 can't be changed; you have to pi=
ck <i>some</i> type to use to handle the storage of a string. And many user=
s will pick std::string because... it's there. Even if they picked the stri=
ng type you like, that may not be the string type other people like. And tr=
ying to design an "omni_string", which is all things to all people, is a tr=
ail that ends in tears.<br><br>#3 generally can't be changed either. Most o=
f such functions require the insertion or deletions of characters. Even see=
mingly simple modifications like upper/lowercase changes are impossible to =
do in Unicode without being able to insert/delete characters. Even in UTF-3=
2, there are codepoint sequences which are considered one case, that map to=
 a single codepoint in another case.<br><br>I'd say the 90% interoperabilit=
y that std::string_view gives you is the best you can reasonably expect to =
get. Take it and move on.<br><br>It should also be noted that your attempt =
at a bunch of policies for std::basic_string wouldn't work for interop anyw=
ay. At least, not without copying. If the library uses a 16-character SSO s=
td::string, but you use a 32-character one, you have to copy when crossing =
that boundary. That's just as true as if it were using a different allocato=
r on its std::string.<br>&nbsp;</div><blockquote class=3D"gmail_quote" styl=
e=3D"margin: 0;margin-left: 0.8ex;border-left: 1px #ccc solid;padding-left:=
 1ex;"><div dir=3D"ltr"><div></div><div>This is something I've personally r=
un into a lot with games, where some otherwise very handy library ends up d=
epending on a verboten STL type, or Boost, or one of the oft verboten C++ f=
eatures, or ends up working around the same problems we have by implementin=
g their own non-standard framework library (and then we'd have to get all t=
hese libraries interoperating with us and each other), and so we can't use =
the library and have to reimplement it from scratch as well as the standard=
 types.</div><div><br></div><div>It's a domino effect of wasted time and mo=
ney.</div></div></blockquote><div><br>This is <i>your</i> problem, not C++'=
s.<br><br>Your problem is that you aren't writing C++ code; you're writing =
to some !C++ language, which is a functional subset of C++. That's all well=
 and good, but the rest of us <i>don't</i> live in your little box. And if =
you expect us to do so, you'll be waiting a while.<br><br>I understand that=
 you didn't necessarily choose your programming environment. But at the end=
 of the day, you can't complain that some C++ library's writer dared to act=
ually <i>use C++</i> in their code, rather than the arbitrary subset you're=
 limited to.<br><br>If you are allergic to Boost, that's something I can un=
derstand. But by rejecting it, you are rejecting every library that uses it=
 as a dependency. And that's a lot of code. That was the choice you made wh=
en you rejected it, and you should have considered that when making that ch=
oice (assuming the choice wasn't made for you). The same goes for any other=
 "oft verboten C++ features" (FYI: Just because something is 'oft' for you =
doesn't make it 'oft' for everyone else too).<br><br>People who actually us=
e the language in their code is not a problem that the language needs to so=
lve. It's only "a domino effect of wasted time and money" for those who hav=
e arbitrary limits placed on what parts of the language they can use. And w=
hile I empathize with you, complaining about it or requesting features to "=
correct" this "problem" is the wrong way to go about it.<br></div></div>

<p></p>

-- <br />
<br />
--- <br />
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br />
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to <a href=3D"mailto:std-proposals+unsubscribe@isocpp.org">std-proposa=
ls+unsubscribe@isocpp.org</a>.<br />
To post to this group, send email to <a href=3D"mailto:std-proposals@isocpp=
..org">std-proposals@isocpp.org</a>.<br />
Visit this group at <a href=3D"http://groups.google.com/a/isocpp.org/group/=
std-proposals/">http://groups.google.com/a/isocpp.org/group/std-proposals/<=
/a>.<br />

------=_Part_1120_1337897824.1433660245863--
------=_Part_1119_157879467.1433660245863--

.
