220 30980 <2eab187f-e1cd-4c7d-ad30-ac7981963749@isocpp.org> article
Path: news.gmane.org!.POSTED!not-for-mail
From: Nicol Bolas <jmckesson@gmail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Proposed alternative approach to specifying
 required memory operation ordering
Date: Fri, 17 Feb 2017 21:40:23 -0800 (PST)
Lines: 445
Approved: news@gmane.org
Message-ID: <2eab187f-e1cd-4c7d-ad30-ac7981963749@isocpp.org>
References: <d0fdabcd-f04a-46e6-b93a-0b9ea85cb5bd@isocpp.org>
 <20170216004134.4902991.82070.24891@gmail.com>
 <8575c41d-1c8e-4e6f-9f12-a07bb70a7703@isocpp.org>
 <a17092d7-d348-4d8c-94f1-fa272ef9e630@isocpp.org>
 <9f578372-9d47-4f5b-800d-93e08e7d2314@isocpp.org>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: blaine.gmane.org
Mime-Version: 1.0
Content-Type: multipart/mixed; 
	boundary="----=_Part_436_2030082738.1487396423706"
X-Trace: blaine.gmane.org 1487396430 31718 195.159.176.226 (18 Feb 2017 05:40:30 GMT)
X-Complaints-To: usenet@blaine.gmane.org
NNTP-Posting-Date: Sat, 18 Feb 2017 05:40:30 +0000 (UTC)
Cc: inkwizytoryankes@gmail.com
To: ISO C++ Standard - Future Proposals <std-proposals@isocpp.org>
Original-X-From: std-proposals+bncBCEKFTV6ZUMBBSF4T7CQKGQE2RIZJBY@isocpp.org Sat Feb 18 06:40:24 2017
Return-path: <std-proposals+bncBCEKFTV6ZUMBBSF4T7CQKGQE2RIZJBY@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-pf0-f197.google.com ([209.85.192.197])
	by blaine.gmane.org with esmtp (Exim 4.84_2)
	(envelope-from <std-proposals+bncBCEKFTV6ZUMBBSF4T7CQKGQE2RIZJBY@isocpp.org>)
	id 1cexkp-0007Kv-T1
	for gclcip-std-proposals@m.gmane.org; Sat, 18 Feb 2017 06:40:20 +0100
Original-Received: by mail-pf0-f197.google.com with SMTP id 204sf88668274pfx.1
        for <gclcip-std-proposals@m.gmane.org>; Fri, 17 Feb 2017 21:40:25 -0800 (PST)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=isocpp-org.20150623.gappssmtp.com; s=20150623;
        h=date:from:to:cc:message-id:in-reply-to:references:subject
         :mime-version:x-original-sender:reply-to:precedence:mailing-list
         :list-id:x-spam-checked-in-group:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe;
        bh=iYg4/f8NGDcnbFZ4x5E/LWUI9nDhmgk6DcgNTqx0kJA=;
        b=gU1qxCvEUR6HDcq9xBjR8jL0eKnwpl2BBVJ5jPNXHvSQ63i54Ks5RbHk/FEZYIY7zR
         5TsqtspDXQl+21g+z6mN7dH1Lyz2W8JNaKt1QKVzd1Wr2uE2hXBUi5gZh+RY/sXSolNV
         DM2LW0Y94tDa5a6SNaiSOUTPbt3kpo2tdIz0GcB7XpqTz8vY8piGlsL32WdFGFyWur+9
         boACuNmEEMMH3OTP29jdp1fo0jsGtMetbJ8qQYPS/ImiwiUuhjSQLKVf1cgBum3yBq7a
         DDxUfugSpbcZWd4HLyOOK/shbmtr36zvRKIFsJF6JfEtF0CXniEls7sL7Xyx6wfrK80s
         4qZg==
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=gmail.com; s=20161025;
        h=date:from:to:cc:message-id:in-reply-to:references:subject
         :mime-version:x-original-sender:reply-to:precedence:mailing-list
         :list-id:x-spam-checked-in-group:list-post:list-help:list-archive
         :list-subscribe:list-unsubscribe;
        bh=iYg4/f8NGDcnbFZ4x5E/LWUI9nDhmgk6DcgNTqx0kJA=;
        b=YeK+G1nElVL66x5PXLDC1iFhVJDgKAyBZD6Xr+cEq+1foi5r5Bc55dj+l+X21Zj/t/
         7bd3BNBxwTdClnny8MYxv9RCVT/J0JR85hyBV/5hjICqB9PhFhH5SOds00t8tdjqULcC
         BkqGqYgGNtnVtWzddudfx7DZBhgJU8ckpbJ2KM9OeI4EXOexyc+RrNd03rMunKS2nvKl
         qkPeTxCwmSjI0+tjuujR0f9yV+bLNTmhdXeWpoJpmTnqJ8EjOAcvTHVR5qcMWb3lAP96
         Up6xdkhZ8Rp6FyQwvjebPr2wnnaRJU6XntW6PZpoymYQeUjLbXC/Y4fjdA4Q3KAuZu89
         WSzw==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20161025;
        h=x-gm-message-state:date:from:to:cc:message-id:in-reply-to
         :references:subject:mime-version:x-original-sender:reply-to
         :precedence:mailing-list:list-id:x-spam-checked-in-group:list-post
         :list-help:list-archive:list-subscribe:list-unsubscribe;
        bh=iYg4/f8NGDcnbFZ4x5E/LWUI9nDhmgk6DcgNTqx0kJA=;
        b=ITdqIv7J3/MX9as8lAa64PXykqnE/ppJ/G/n1iudZQUgBaDIm9hgS4RszSTeSmgaLE
         uTOoOjIwjHGukLwQ1rgyfWKqYcgAuBtdwfnKJFOxeMsuRBJXQfCGCkiNdBXTAm9EY+TD
         uBd3DXZiuIvsz45CBX5HzAaL0fWGhVAUmZLsF1Edu0tf8UVOWiVhT5no65ze5viKf0pd
         owh5ipIBlJZ6kuLjWvIM+3UZW1WQ3o3TGyURV7tisKjxyLzXFaw+K8/zaGOnYgVnvRcv
         qXL7iBIR05Aqk/jTO8omKTedYS92H50SLPcptN0DtF7P8SZ7l1kBacWeeczL3cg/q0+G
         dwLQ==
X-Gm-Message-State: AMke39nNA7tqNGP6XjZzv2FYTYVgtwtFjxeWI98DnNyZAgMCrZgFRdggR0KAh+x+qsDkyw==
X-Received: by 10.98.78.6 with SMTP id c6mr2986157pfb.40.1487396425253;
        Fri, 17 Feb 2017 21:40:25 -0800 (PST)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.157.34.169 with SMTP id y38ls6002415ota.11.gmail; Fri, 17 Feb
 2017 21:40:24 -0800 (PST)
X-Received: by 10.157.62.29 with SMTP id a29mr957778otd.5.1487396424373;
        Fri, 17 Feb 2017 21:40:24 -0800 (PST)
In-Reply-To: <9f578372-9d47-4f5b-800d-93e08e7d2314@isocpp.org>
X-Original-Sender: jmckesson@gmail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <https://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <https://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <https://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <https://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>,
 <https://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:30980
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/30980>

------=_Part_436_2030082738.1487396423706
Content-Type: multipart/alternative; 
	boundary="----=_Part_437_796159868.1487396423706"

------=_Part_437_796159868.1487396423706
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

On Friday, February 17, 2017 at 11:27:57 PM UTC-5, Walt Karas wrote:
>
> On Friday, February 17, 2017 at 5:06:14 PM UTC-5, inkwizyt...@gmail.com=
=20
> wrote:
>>
>> On Friday, February 17, 2017 at 8:54:24 PM UTC+1, Walt Karas wrote:
>>>
>>> On Wednesday, February 15, 2017 at 7:41:40 PM UTC-5, Tony V E wrote:
>>>>
>>>> Leaving the questions about =E2=80=8Eactually changing the standard as=
ide, and=20
>>>> focusing on understanding (which may just mean this could be a=20
>>>> std-discussion question instead of std-proposal),
>>>>
>>>> In your model, if X and Y are relaxed atomic operations in thread T1,=
=20
>>>> can thread T2 see them as Y before X while thread T3 sees them as =E2=
=80=8EX before=20
>>>> Y?
>>>>
>>>> Sent from my BlackBerry portable Babbage Device
>>>> *From: *'Walt Karas' via ISO C++ Standard - Future Proposals
>>>> *Sent: *Wednesday, February 15, 2017 2:17 PM
>>>> *To: *ISO C++ Standard - Future Proposals
>>>> *Reply To: *std-pr...@isocpp.org
>>>> *Subject: *[std-proposals] Proposed alternative approach to specifying=
=20
>>>> required memory operation ordering
>>>>
>>>> - In a program execution, each thread defines a nominal order of=20
>>>> (thread local) memory and fence operations.
>>>> - For any operations X and Y in thread T, either X before Y, or Y=20
>>>> before X.
>>>> - A program execution defines a nominal global order of global memory=
=20
>>>> operations.
>>>> - If X and Y are atomic global operations, then either X before Y or Y=
=20
>>>> before X in the global order.
>>>> - If X is a global operation, and there exists a global store S where=
=20
>>>> neither X before S nor S before X in the global order, then the result=
 of X=20
>>>> is undefined.
>>>> - A program execution defines a partial function f(T, LO) -> GO where=
=20
>>>> LO is a local memory operation in thread T, and GO is a global operati=
on. =20
>>>> If LO is atomic, then GO must be atomic. The result of LO is the resul=
t of=20
>>>> GO.  If the result of GO is undefined, the result of LO is undefined. =
=20
>>>> (Even if f(T, LO) is not required to exist, it none the less _may_ exi=
st.)
>>>> - If LO1 and LO2 are operations in thread T, and LO1 before LO2 in T,=
=20
>>>> and f(T, LO1) and f(T, LO2) both exist, then f(T, LO2) before f(T, LO1=
) in=20
>>>> the global order is not allowed.
>>>> - If:
>>>> 1.  In a thread T, X and Y are memory operations, and F is a fence=20
>>>> operation.
>>>> 2.  X before F and F before Y.
>>>> 3.  F is sequentially consistent and both X and Y are atomic, or
>>>> 4.  F is acquire and both X and Y are loads, or
>>>> 5.  F is release and both X and Y are stores.
>>>> then F is activated for X and Y.
>>>> - In a thread T, if X and Y are memory operations (where X before Y)=
=20
>>>> with an activated fence F, then f(T, X) must exist.
>>>> - For every global operation GO, there must exist a local operation LO=
=20
>>>> in some thread T where f(T, LO) -> GO.  (Assuming no intense gamma=20
>>>> radiation.)
>>>> - A sequentially consistent (thread local) atomic memory operation=20
>>>> implies two sequentially consistent fences, one before and one after i=
t (as=20
>>>> well as a preceding release for a store, and a succeeding acquire for =
a=20
>>>> load).
>>>>
>>>> --=20
>>>> You received this message because you are subscribed to the Google=20
>>>> Groups "ISO C++ Standard - Future Proposals" group.
>>>> To unsubscribe from this group and stop receiving emails from it, send=
=20
>>>> an email to std-proposal...@isocpp.org.
>>>> To post to this group, send email to std-pr...@isocpp.org.
>>>> To view this discussion on the web visit=20
>>>> https://groups.google.com/a/isocpp.org/d/msgid/std-proposals/d0fdabcd-=
f04a-46e6-b93a-0b9ea85cb5bd%40isocpp.org=20
>>>> <https://groups.google.com/a/isocpp.org/d/msgid/std-proposals/d0fdabcd=
-f04a-46e6-b93a-0b9ea85cb5bd%40isocpp.org?utm_medium=3Demail&utm_source=3Df=
ooter>
>>>> .
>>>>
>>>>
>>> Here is some code that illustrates some points of confusion I have with=
=20
>>> the memory model:
>>>
>>> (...)
>>>
>>>         while (*Twins_number < equals_mine)
>>>           ;
>>>
>> =20
>> I'm not expert or even experienced in atomics but I think that is race=
=20
>> condition because you access this variable without synchronization. You=
=20
>> would need probably fence or acquire in `while` otherwise compiler can=
=20
>> amuse that value will not change there.
>>
>
> I don't really understand that.  To me, the best way to describe the=20
> problem is that *Twins_number is cached in a (thread-specific) register. =
=20
> By some mechanism, a cache-invalidate or equivalent has to be performed o=
n=20
> *Twins_number before reading it.
>

That's because you're thinking in terms of how it works on the actual=20
hardware, not in terms of the abstract C++ memory model. The *whole point*=
=20
of an abstract memory model is so that the C++ coder doesn't have to know=
=20
or care about which "actual hardware" their code runs on. If you follow the=
=20
rules of the memory model, the compiler for the "actual hardware" will=20
generate whatever code is needed to make your stuff work.
=20

> The other issue is that, using gcc, this works when the read of=20
> *Twins_number is atomic with relaxed order.  I think that the Standard=20
> allows for this to be just a broken as when using a (actual or nominal)=
=20
> non-atomic read, and it's just a qwerky implementation in gcc.
>

Actually... no, it isn't.

I was mistaken earlier about something. The memory orders=20
<http://en.cppreference.com/w/cpp/atomic/memory_order> are not for ensuring=
=20
visibility of atomic variables across threads. Atomic variables are=20
*always* visible across threads. Their memory orders are for ensuring=20
ordering and visibility for *non-atomic* stuff which is done around some=20
atomic operation.

For example, let's say you write some data on thread A, then set an atomic=
=20
boolean to true. You have a thread B that wants to read that data. It can=
=20
spin-lock on that boolean and read the data once the boolean is true.

However, the standard will *only* guarantee this if you use proper ordering=
=20
constraints. Thread A *must* have used a `store` operation on that boolean=
=20
with `release` ordering or stronger. And thread B *must* have used a `load`=
=20
with `acquire` ordering or stronger in its spin-lock. Doing anything less=
=20
yields a data race and therefore UB.

But if all you're using an atomic for is the value itself (note: this does=
=20
not apply to the value of the object it points to, for atomic pointers),=20
then `relaxed` memory ordering is just fine. So long as you don't care=20
about the ordering of independent atomic operations on different variables.
=20

> Maybe using Cardinal =3D unsigned would work with optimization enabled if=
 I=20
> put an acquire fence in the loop.  But that seems a cheat.
>

Why is that a cheat? You need to tell the compiler that operations from=20
another thread may affect the execution with this one. That's *exactly*=20
what a fence does.

And since you don't know exactly when the writing thread executes its=20
`release` fence, you must execute your `acquire` fence in the loop.

Unless the memory model considers that all the loads of *Twins_number can=
=20
> happen simultaneously, and the acquire fence is the correct way to cause=
=20
> them to happen sequentially.
>

Remember: a data race is undefined behavior in C++. Therefore, if I have a=
=20
`while(*some_ptr);` infinite loop, the compiler is *completely free* to=20
convert this into `if(*some_ptr) while(true);`. Why?

Because a data race is undefined behavior. There is nothing in your=20
infinite loop that would ensure visibility between this thread and another=
=20
thread. And therefore, if some other thread modified the memory at=20
`some_ptr`, your thread reading it without proper synchronization would be=
=20
a data race and therefore UB.

So, if some other thread modifies the memory of `*some_ptr`, then you get=
=20
UB. And therefore, the compiler is free to assume that your code actually=
=20
works, and therefore no other thread modifies that pointer. So if no other=
=20
thread modifies that pointer's value, then `if(*some_ptr) while(true);` is=
=20
a perfectly justifiable interpretation of your code.

And if the compiler can prove that the condition will *never* be false=20
(since the only way it could be false would be for another thread to modify=
=20
the memory, which would be UB), then it can just reduce the code to a pure=
=20
infinite loop.

--=20
You received this message because you are subscribed to the Google Groups "=
ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
To view this discussion on the web visit https://groups.google.com/a/isocpp=
..org/d/msgid/std-proposals/2eab187f-e1cd-4c7d-ad30-ac7981963749%40isocpp.or=
g.

------=_Part_437_796159868.1487396423706
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">On Friday, February 17, 2017 at 11:27:57 PM UTC-5, Walt Ka=
ras wrote:<blockquote class=3D"gmail_quote" style=3D"margin: 0;margin-left:=
 0.8ex;border-left: 1px #ccc solid;padding-left: 1ex;"><div dir=3D"ltr">On =
Friday, February 17, 2017 at 5:06:14 PM UTC-5, <a>inkwizyt...@gmail.com</a>=
 wrote:<blockquote class=3D"gmail_quote" style=3D"margin:0;margin-left:0.8e=
x;border-left:1px #ccc solid;padding-left:1ex"><div dir=3D"ltr">On Friday, =
February 17, 2017 at 8:54:24 PM UTC+1, Walt Karas wrote:<blockquote class=
=3D"gmail_quote" style=3D"margin:0;margin-left:0.8ex;border-left:1px #ccc s=
olid;padding-left:1ex"><div dir=3D"ltr">On Wednesday, February 15, 2017 at =
7:41:40 PM UTC-5, Tony V E wrote:<blockquote class=3D"gmail_quote" style=3D=
"margin:0;margin-left:0.8ex;border-left:1px #ccc solid;padding-left:1ex"><d=
iv style=3D"background-color:rgb(255,255,255);line-height:initial" lang=3D"=
en-US">                                                                    =
                  <div style=3D"width:100%;font-size:initial;font-family:Ca=
libri,&#39;Slate Pro&#39;,sans-serif,sans-serif;color:rgb(31,73,125);text-a=
lign:initial;background-color:rgb(255,255,255)">Leaving the questions about=
 =E2=80=8Eactually changing the standard aside, and focusing on understandi=
ng (which may just mean this could be a std-discussion question instead of =
std-proposal),</div><div style=3D"width:100%;font-size:initial;font-family:=
Calibri,&#39;Slate Pro&#39;,sans-serif,sans-serif;color:rgb(31,73,125);text=
-align:initial;background-color:rgb(255,255,255)"><br></div><div style=3D"w=
idth:100%;font-size:initial;font-family:Calibri,&#39;Slate Pro&#39;,sans-se=
rif,sans-serif;color:rgb(31,73,125);text-align:initial;background-color:rgb=
(255,255,255)">In your model, if X and Y are relaxed atomic operations in t=
hread T1, can thread T2 see them as Y before X while thread T3 sees them as=
 =E2=80=8EX before Y?</div>                                                =
                                                                           =
          <div style=3D"width:100%;font-size:initial;font-family:Calibri,&#=
39;Slate Pro&#39;,sans-serif,sans-serif;color:rgb(31,73,125);text-align:ini=
tial;background-color:rgb(255,255,255)"><br style=3D"display:initial"></div=
>                                                                          =
                                                                           =
                                              <div style=3D"font-size:initi=
al;font-family:Calibri,&#39;Slate Pro&#39;,sans-serif,sans-serif;color:rgb(=
31,73,125);text-align:initial;background-color:rgb(255,255,255)">Sent=C2=A0=
from=C2=A0my=C2=A0BlackBerry=C2=A0<wbr>portable=C2=A0Babbage=C2=A0Device</d=
iv>                                                                        =
                                                                           =
                               <table style=3D"background-color:white;borde=
r-spacing:0px" width=3D"100%"> <tbody><tr><td style=3D"font-size:initial;te=
xt-align:initial;background-color:rgb(255,255,255)" colspan=3D"2">         =
                  <div style=3D"border-style:solid none none;border-top-col=
or:rgb(181,196,223);border-top-width:1pt;padding:3pt 0in 0in;font-family:Ta=
homa,&#39;BB Alpha Sans&#39;,&#39;Slate Pro&#39;;font-size:10pt">  <div><b>=
From: </b>&#39;Walt Karas&#39; via ISO C++ Standard - Future Proposals</div=
><div><b>Sent: </b>Wednesday, February 15, 2017 2:17 PM</div><div><b>To: </=
b>ISO C++ Standard - Future Proposals</div><div><b>Reply To: </b><a rel=3D"=
nofollow">std-pr...@isocpp.org</a></div><div><b>Subject: </b>[std-proposals=
] Proposed alternative approach to specifying required memory operation ord=
ering</div></div></td></tr></tbody></table><div style=3D"border-style:solid=
 none none;border-top-color:rgb(186,188,209);border-top-width:1pt;font-size=
:initial;text-align:initial;background-color:rgb(255,255,255)"></div><br><d=
iv><div dir=3D"ltr"><div>- In a program execution,=C2=A0each thread defines=
 a nominal order of (thread local) memory and fence operations.</div><div>-=
 For any operations X and Y in thread T, either X before Y, or Y before X.<=
/div><div>- A program execution defines a nominal global order of global me=
mory operations.</div><div>- If X and Y are atomic global operations, then =
either X before Y or Y before X in the global order.</div><div>- If X is a =
global operation, and there exists a global store S where neither X before =
S nor S before X in the global order, then the result of X is undefined.</d=
iv><div>- A program execution defines a partial function f(T, LO) -&gt; GO =
where LO is a local memory operation in thread T, and GO is a global operat=
ion.=C2=A0 If LO is atomic, then GO must be atomic. The result of LO is the=
 result of GO.=C2=A0 If the result of GO is undefined, the result of LO is =
undefined.=C2=A0 (Even if f(T, LO) is not required to exist, it none the le=
ss _may_ exist.)</div><div>- If LO1 and LO2 are operations in thread T, and=
 LO1 before LO2 in T, and f(T, LO1) and f(T, LO2) both exist, then f(T, LO2=
) before f(T, LO1) in the global order is not allowed.</div><div>- If:</div=
><div>1.=C2=A0 In a thread T, X and Y are memory operations, and F is a fen=
ce operation.</div><div>2.=C2=A0 X before=C2=A0F and=C2=A0F before Y.</div>=
<div>3.=C2=A0 F is sequentially consistent and both X and Y are atomic, or<=
/div><div>4.=C2=A0 F is acquire and both X and Y are loads, or</div><div>5.=
=C2=A0 F is release and both X and Y are stores.</div><div>then F is activa=
ted for X and Y.</div><div>-=C2=A0In a thread T, if X and Y are memory oper=
ations (where X before Y) with an activated fence F, then f(T, X) must exis=
t.</div><div>- For every global operation GO, there must exist a local oper=
ation LO in some thread T where f(T, LO) -&gt; GO.=C2=A0 (Assuming no inten=
se gamma radiation.)</div><div>- A sequentially consistent (thread local) a=
tomic memory operation implies two sequentially consistent fences, one befo=
re and one after it (as well as a preceding release for a store, and a succ=
eeding acquire for a load).</div></div>

<p></p>

-- <br>
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br>
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to <a rel=3D"nofollow">std-proposal...@isocpp.org</a>.<br>
To post to this group, send email to <a rel=3D"nofollow">std-pr...@isocpp.o=
rg</a>.<br>
To view this discussion on the web visit <a href=3D"https://groups.google.c=
om/a/isocpp.org/d/msgid/std-proposals/d0fdabcd-f04a-46e6-b93a-0b9ea85cb5bd%=
40isocpp.org?utm_medium=3Demail&amp;utm_source=3Dfooter" rel=3D"nofollow" t=
arget=3D"_blank" onmousedown=3D"this.href=3D&#39;https://groups.google.com/=
a/isocpp.org/d/msgid/std-proposals/d0fdabcd-f04a-46e6-b93a-0b9ea85cb5bd%40i=
socpp.org?utm_medium\x3demail\x26utm_source\x3dfooter&#39;;return true;" on=
click=3D"this.href=3D&#39;https://groups.google.com/a/isocpp.org/d/msgid/st=
d-proposals/d0fdabcd-f04a-46e6-b93a-0b9ea85cb5bd%40isocpp.org?utm_medium\x3=
demail\x26utm_source\x3dfooter&#39;;return true;">https://groups.google.com=
/a/<wbr>isocpp.org/d/msgid/std-<wbr>proposals/d0fdabcd-f04a-46e6-<wbr>b93a-=
0b9ea85cb5bd%40isocpp.org</a><wbr>.<br>
<br></div></div></blockquote><div><br></div><div>Here is some code that ill=
ustrates some points of confusion I have with the memory model:</div><div><=
br></div><div><font color=3D"#666600">(...)</font><br><div><font face=3D"mo=
nospace" color=3D"#666600"><br></font></div><div><font face=3D"monospace" c=
olor=3D"#666600">=C2=A0 =C2=A0 =C2=A0 =C2=A0 while (*Twins_number &lt; equa=
ls_mine)</font></div><div><font face=3D"monospace" color=3D"#666600">=C2=A0=
 =C2=A0 =C2=A0 =C2=A0 =C2=A0 ;</font></div></div></div></blockquote><div>=
=C2=A0<br>I&#39;m not expert or even experienced in atomics but I think tha=
t is race condition because you access this variable without synchronizatio=
n. You would need probably fence or acquire in `while` otherwise compiler c=
an amuse that value will not change there.<br></div></div></blockquote><div=
><br></div><div>I don&#39;t=C2=A0really understand that.=C2=A0 To me, the b=
est way to describe the problem is that *Twins_number is cached in a (threa=
d-specific) register.=C2=A0 By some mechanism, a cache-invalidate or equiva=
lent has=C2=A0to be performed on *Twins_number before reading it.</div></di=
v></blockquote><div><br>That&#39;s because you&#39;re thinking in terms of =
how it works on the actual hardware, not in terms of the abstract C++ memor=
y model. The <i>whole point</i> of an abstract memory model is so that the =
C++ coder doesn&#39;t have to know or care about which &quot;actual hardwar=
e&quot; their code runs on. If you follow the rules of the memory model, th=
e compiler for the &quot;actual hardware&quot; will generate whatever code =
is needed to make your stuff work.<br>=C2=A0</div><blockquote class=3D"gmai=
l_quote" style=3D"margin: 0;margin-left: 0.8ex;border-left: 1px #ccc solid;=
padding-left: 1ex;"><div dir=3D"ltr"><div></div><div>The other issue is tha=
t, using gcc, this works when the read of *Twins_number is=C2=A0atomic=C2=
=A0with relaxed order.=C2=A0 I think that the Standard allows for this to b=
e just a broken as when using a=C2=A0(actual or nominal) non-atomic read, a=
nd it&#39;s just a qwerky implementation in gcc.</div></div></blockquote><d=
iv><br>Actually... no, it isn&#39;t.<br><br>I was mistaken earlier about so=
mething. The <a href=3D"http://en.cppreference.com/w/cpp/atomic/memory_orde=
r">memory orders</a> are not for ensuring visibility of atomic variables ac=
ross threads. Atomic variables are *always* visible across threads. Their m=
emory orders are for ensuring ordering and visibility for <i>non-atomic</i>=
 stuff which is done around some atomic operation.<br><br>For example, let&=
#39;s say you write some data on thread A, then set an atomic boolean to tr=
ue. You have a thread B that wants to read that data. It can spin-lock on t=
hat boolean and read the data once the boolean is true.<br><br>However, the=
 standard will <i>only</i> guarantee this if you use proper ordering constr=
aints. Thread A <i>must</i> have used a `store` operation on that boolean w=
ith `release` ordering or stronger. And thread B <i>must</i> have used a `l=
oad` with `acquire` ordering or stronger in its spin-lock. Doing anything l=
ess yields a data race and therefore UB.<br><br>But if all you&#39;re using=
 an atomic for is the value itself (note: this does not apply to the value =
of the object it points to, for atomic pointers), then `relaxed` memory ord=
ering is just fine. So long as you don&#39;t care about the ordering of ind=
ependent atomic operations on different variables.<br>=C2=A0</div><blockquo=
te class=3D"gmail_quote" style=3D"margin: 0;margin-left: 0.8ex;border-left:=
 1px #ccc solid;padding-left: 1ex;"><div dir=3D"ltr"><div>Maybe using Cardi=
nal =3D unsigned would work with optimization enabled if I put an acquire f=
ence in the loop.=C2=A0 But that seems a cheat.</div></div></blockquote><di=
v><br>Why is that a cheat? You need to tell the compiler that operations fr=
om another thread may affect the execution with this one. That&#39;s <i>exa=
ctly</i> what a fence does.<br><br>And since you don&#39;t know exactly whe=
n the writing thread executes its `release` fence, you must execute your `a=
cquire` fence in the loop.<br><br></div><blockquote class=3D"gmail_quote" s=
tyle=3D"margin: 0;margin-left: 0.8ex;border-left: 1px #ccc solid;padding-le=
ft: 1ex;"><div dir=3D"ltr"><div>Unless the memory model considers that all =
the loads of *Twins_number can happen simultaneously, and the acquire fence=
 is the correct way to cause them to happen sequentially.</div></div></bloc=
kquote><div><br>Remember: a data race is undefined behavior in C++. Therefo=
re, if I have a `while(*some_ptr);` infinite loop, the compiler is <i>compl=
etely free</i> to convert this into `if(*some_ptr) while(true);`. Why?<br><=
br>Because a data race is undefined behavior. There is nothing in your infi=
nite loop that would ensure visibility between this thread and another thre=
ad. And therefore, if some other thread modified the memory at `some_ptr`, =
your thread reading it without proper synchronization would be a data race =
and therefore UB.<br><br>So, if some other thread modifies the memory of `*=
some_ptr`, then you get UB. And therefore, the compiler is free to assume t=
hat your code actually works, and therefore no other thread modifies that p=
ointer. So if no other thread modifies that pointer&#39;s value, then `if(*=
some_ptr) while(true);` is a perfectly justifiable interpretation of your c=
ode.<br><br>And if the compiler can prove that the condition will <i>never<=
/i> be false (since the only way it could be false would be for another thr=
ead to modify the memory, which would be UB), then it can just reduce the c=
ode to a pure infinite loop.<br></div></div>

<p></p>

-- <br />
You received this message because you are subscribed to the Google Groups &=
quot;ISO C++ Standard - Future Proposals&quot; group.<br />
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to <a href=3D"mailto:std-proposals+unsubscribe@isocpp.org">std-proposa=
ls+unsubscribe@isocpp.org</a>.<br />
To post to this group, send email to <a href=3D"mailto:std-proposals@isocpp=
..org">std-proposals@isocpp.org</a>.<br />
To view this discussion on the web visit <a href=3D"https://groups.google.c=
om/a/isocpp.org/d/msgid/std-proposals/2eab187f-e1cd-4c7d-ad30-ac7981963749%=
40isocpp.org?utm_medium=3Demail&utm_source=3Dfooter">https://groups.google.=
com/a/isocpp.org/d/msgid/std-proposals/2eab187f-e1cd-4c7d-ad30-ac7981963749=
%40isocpp.org</a>.<br />

------=_Part_437_796159868.1487396423706--

------=_Part_436_2030082738.1487396423706--

.
