220 12571 <53FFA90E.4080501@gmx.net> article
Path: news.gmane.org!not-for-mail
From: Jens Maurer <Jens.Maurer@gmx.net>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Cryptographic hash functions reloaded [was
 Interest in cryptographic functions within the standard library]
Date: Fri, 29 Aug 2014 00:11:26 +0200
Lines: 189
Approved: news@gmane.org
Message-ID: <53FFA90E.4080501@gmx.net>
References: <53F8F620.3090701@gmx.de> <B33BCE97-51A4-4880-9C26-5347D6C8A02A@gmail.com> <53F989E9.5090701@gmx.net> <B0F486D7-A068-4A30-8D5E-69E826F1C943@gmail.com> <53FE5E8A.2090003@gmx.net> <6B0F9092-02BD-46A1-871F-14746F3938BA@gmail.com>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable
X-Trace: ger.gmane.org 1409263896 16267 80.91.229.3 (28 Aug 2014 22:11:36 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Thu, 28 Aug 2014 22:11:36 +0000 (UTC)
To: std-proposals@isocpp.org
Original-X-From: std-proposals+bncBDPMTYGK64JRBEOS72PQKGQEWZVK4EY@isocpp.org Fri Aug 29 00:11:31 2014
Return-path: <std-proposals+bncBDPMTYGK64JRBEOS72PQKGQEWZVK4EY@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-ee0-f72.google.com ([74.125.83.72])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBDPMTYGK64JRBEOS72PQKGQEWZVK4EY@isocpp.org>)
	id 1XN7uk-0008Oc-FE
	for gclcip-std-proposals@m.gmane.org; Fri, 29 Aug 2014 00:11:30 +0200
Original-Received: by mail-ee0-f72.google.com with SMTP id e49sf1512328eek.7
        for <gclcip-std-proposals@m.gmane.org>; Thu, 28 Aug 2014 15:11:30 -0700 (PDT)
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:message-id:date:from:user-agent:mime-version:to
         :subject:references:in-reply-to:x-original-sender
         :x-original-authentication-results:reply-to:precedence:mailing-list
         :list-id:list-post:list-help:list-archive:list-subscribe
         :list-unsubscribe:content-type:content-transfer-encoding;
        bh=2olQ/9gOhBW3z0THoCNQkYOzYXZja4GVuJ7CpqI02OU=;
        b=W0L40rdbgl6KuaNGskEbXxWi4yXeA0MGMngejQrVLoEWZoUJesfzh7irIywnMhySzZ
         sfRsPSMQTyJWzcp216kcq1dH2Aa1Hv/fTb7oZDTlAQV5El4GFCyzPejSnmR2W6tvn4sG
         zyrwJGNT6I5NVrP65+afRHba14Bi/pR9aSH9mcH9iFVVA0Pl3nQ6T9kfwrHKCfLkalbH
         eH1hm7tUUFdLepUNXFe2HKh2npO4KZJMeeOuWzTdIKLY+VbCSL4W3LOKiKFTpdfEDx2D
         m8vm0Ljp+PoSNVj1UjwD5xCpXERmRYcJr11gywetIqF5Nmor43juAPhyEeMWA2+tyJBn
         FUc 
X-Gm-Message-State: ALoCoQmGXuDG26pANXDL3AbAzlUYythW6WIbkO/10f8R0w5+QBdKRMsDmsK7PQ0OQLdJqS5122+P
X-Received: by 10.180.36.98 with SMTP id p2mr693547wij.0.1409263890187;
        Thu, 28 Aug 2014 15:11:30 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.152.45.9 with SMTP id i9ls211604lam.99.gmail; Thu, 28 Aug 2014
 15:11:28 -0700 (PDT)
X-Received: by 10.152.234.162 with SMTP id uf2mr6165782lac.44.1409263888569;
        Thu, 28 Aug 2014 15:11:28 -0700 (PDT)
Original-Received: from mout.gmx.net (mout.gmx.net. [212.227.15.19])
        by mx.google.com with ESMTPS id u3si7276838lbv.132.2014.08.28.15.11.28
        for <std-proposals@isocpp.org>
        (version=TLSv1.2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128);
        Thu, 28 Aug 2014 15:11:28 -0700 (PDT)
Received-SPF: pass (google.com: domain of Jens.Maurer@gmx.net designates 212.227.15.19 as permitted sender) client-ip=212.227.15.19;
Original-Received: from [192.168.2.100] ([178.7.39.200]) by mail.gmx.com (mrgmx001)
 with ESMTPSA (Nemesis) id 0MDQp3-1X6KXa2fyX-00GoF1 for
 <std-proposals@isocpp.org>; Fri, 29 Aug 2014 00:11:27 +0200
User-Agent: Mozilla/5.0 (X11; Linux i686; rv:6.0.1) Gecko/20110830 Thunderbird/6.0.1
In-Reply-To: <6B0F9092-02BD-46A1-871F-14746F3938BA@gmail.com>
X-Provags-ID: V03:K0:uvuc7vk4PFcBw7wKX+IvNYXfsnsNPAforfmwJLqai/+evD4vdLM
 4ptBYxza6Px1qM7QAG/tWw6g3XhNxa6qdzPK0Ra5yV8TMcOIs+saMn7VggCF4CBHWxTYtE2
 3OQRxbi/PZnnvjBZ0FC1GTr1wEu7tdsDjKrd7BmeniYNwE5VPXo51F8D/oM47OB61ggd9D+
 iFc9uL8AP57OcivohntFg==
X-UI-Out-Filterresults: notjunk:1;
X-Original-Sender: Jens.Maurer@gmx.net
X-Original-Authentication-Results: mx.google.com;       spf=pass (google.com:
 domain of Jens.Maurer@gmx.net designates 212.227.15.19 as permitted sender) smtp.mail=Jens.Maurer@gmx.net
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>,
 <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:12571
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/12571>

On 08/28/2014 07:41 PM, Howard Hinnant wrote:
> I suspect you intended this as a private message, but it came through pub=
lic.=20

It was intended as a public message.  Let's try something different:

Hi everybody!

>> I'll have to ask a few more questions here.  If something gets
>> standardized in this area, I'd like to see a roadmap how all
>> platforms supported by C++ can get portable (crypto) hash values,
>> even if not maximally efficient on "strange" environments.
>>
>> I'm assuming that (in an abstract sense) a crypto hash algorithm
>> (and most others) hash a sequence of octets (i.e. 8-bit quantities).
>>
>> (SHA256 and others can do odd trailing bits, but let's ignore
>> this for now.)
>>
>> I presume I'm passing these octets to the hash algorithm using an
>> array of (unsigned) char, right?
>=20
> As currently coded, there are hash_append overloads for both unsigned cha=
r, C-arrays of unsigned char, std::array<unsigned char, N>, etc.  There are=
 also hash_append overloads for all other arithmetic types, and std-defined=
 containers.  For each type, the hash_append function is responsible for de=
ciding how that type should present itself to a generic hashing algorithm.

I agree that the standard library should provide hash_append overloads for
all scalar types and standard containers, including C-style arrays.
(Btw, does hash_append on a container also hash the size, or just the
contents, i.e. the sequence of elements?)

That's not what I'm concerned about here.  I'm concerned about writing a
crypto hash sum that plays nicely with the framework, and is maximally
portable both in its implementation and its result value.

> For example an unsigned char would just say: Consume this byte!

A byte is not (necessarily) an octet.  I'm concerned about this
particular gap.

>> Is this assumption also true for platforms where (unsigned) char
>> is e.g. 32-bits (DSPs)?  If so, the following implementation doesn't
>> work there, because it makes no effort to split e.g. T=3D=3Dint (suppose
>> it's 32-bit) into four individual (unsigned) char objects.
>=20
> The hash_append infrastructure makes no assumptions on the size of a byte=
..  However concrete hashing algorithms (such as SHA256) most certainly will=
 make such assumptions.

I want to write a portable implementation of a hash sum that also
works on a machine where 1 =3D=3D sizeof(char) =3D=3D sizeof(int) (=3D 32 b=
its).

If I've understood the interface correctly, both hashing an int x and
a char c will end up with a call to my hash algorithm h like this:

   h(&x, 1);
   h(&c, 1);

and the interface is type-erased (i.e. uses "void*" or "unsigned char*").

On a usual platform, the calls will end up like this (ignoring endianess
for now):

   h(&x, 4);
   h(&c, 1);

I don't think "h" will ever be able to produce the same hash sum on both
platforms, even if specifically tailored for the particular platform.
It seems too much information is lost on the   sizeof(char) =3D=3D sizeof(i=
nt)
platform.

One way to address this is to split the "int" into four octets and
assign a separate "unsigned char" for each octet in the hash_append
function on the DSP-style platform.  Then both calls end up as

   h(&x, 4);
   h(&c, 1);


>> On 08/24/2014 07:39 PM, Howard Hinnant wrote:
>>> If we are dealing with a platform/HashAlgorithm disagreement in endian,=
 then an alternative hash_append can be used for scalars:
>>
>> So, the endianness is a boolean, not a three-way type?  Either you're "n=
ative" or not
>> seems all that matters, from the code you presented.
>=20
> As currently coded, a hashing algorithm would set a member static const e=
num to one of three values:
>=20
>     static constexpr xstd::endian endian =3D xstd::endian::native;
>     static constexpr xstd::endian endian =3D xstd::endian::big;
>     static constexpr xstd::endian endian =3D xstd::endian::little;
>=20
> The meaning of these is to ask the hash_append overload for scalars influ=
enced by endian (larger than char) to change the endian from native, to the=
 requested endian, prior to sending the bytes into the hashing algorithm.  =
Concretely, in the order shown above:
>=20
> 1.  Map native endian to native endian (presumably this is always a no-op=
).
> 2.  Map native endian to big endian.  This will be a no-op on big endian =
machines.
> 3.  Map native endian to little endian.  This will be a no-op on little e=
ndian machines.

> Non-fingerprinting hash applications will probably always use the native =
mapping (i.e. they don't care about endian).

Yes.

It seems to me that these choices are, strictly speaking, not a property of=
 the
(crypto) hash algorithm (that is only concerned with octets coming in), but=
 with
the preferences / situation in which it is used.  As someone else pointed o=
ut,
we're essentially defining an ephemeral serialization format for purposes o=
f
computing the hash.

I'd like to ask:

 - that (core) hash algorithm implementations such as SHA256 do not
specify the "endian" thing (it doesn't mean anything at this level) and

 - that there is a config option to do scalar endian conversions if so desi=
red.

Example:

  std::uhash<std::sha256>        // unportable for scalars > char
  std::uhash<std::sha256, std::endian::big>   // convert scalars > char to =
"big endian" prior to feeding octets to hash algorithm

This also supports strange VAX-endianess as the native endian
convention, I believe.

>> The other aspect is the fact that hash algorithms such as sha256 like
>> to process e.g. 32-bits (=3D 4 octets) at once.  When reading four
>> octets from memory, it's helpful to be able to simply read them
>> into a register on "suitable" platforms and only do the endianess
>> (or other) conversion on the remainder of the platforms.  But, on
>> the abstract level, this is not a configuration option, it's a
>> question of correctness.

> Agreed.  I see the second aspect above is an implementation detail of the=
 hashing algorithm.  This detail does not impact the hashing algorithm's in=
terface, except as to impact its results.  I see no motivation to "leak" th=
is implementation detail into a standard specification, unless the standard=
 is to specify concrete hashing algorithms.  In that event we could choose =
any number of options, of which I really have little opinion.  For example:=
  Sha256_output_little_endian as one hashing algorithm and Sha256_output_bi=
g_endian as another.  Or perhaps Sha256<endian> is another solution. =20

Well, there is a lost optimization opportunity if you're hashing an
array of unsigned int (32 bits) on a platform and with an endianess
choice that is just "right".  Otherwise, you get two endianess conversions
back-to-back: One for the scalar > char thing from above, and one when
sha256 tries to form its internal 32-bit chunks.

>> The simple name must result in the portable hash value.

Let me retract that; see details above.

> My understanding from http://en.wikipedia.org/wiki/Sha256 is that SHA256 =
output endian is always big.

The output is a sequence of octets.  It might be that this sequence is inte=
rpreted in
big endian style for (some) display purposes, but that's a minor detail.

Jens

--=20

---=20
You received this message because you are subscribed to the Google Groups "=
ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposa=
ls/.

.
