220 12464 <B0F486D7-A068-4A30-8D5E-69E826F1C943@gmail.com> article
Path: news.gmane.org!not-for-mail
From: Howard Hinnant <howard.hinnant@gmail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Cryptographic hash functions reloaded [was
 Interest in cryptographic functions within the standard library]
Date: Sun, 24 Aug 2014 13:39:53 -0400
Lines: 173
Approved: news@gmane.org
Message-ID: <B0F486D7-A068-4A30-8D5E-69E826F1C943@gmail.com>
References: <53F8F620.3090701@gmx.de> <B33BCE97-51A4-4880-9C26-5347D6C8A02A@gmail.com> <53F989E9.5090701@gmx.net>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0 (Mac OS X Mail 7.3 \(1878.6\))
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable
X-Trace: ger.gmane.org 1408902005 32395 80.91.229.3 (24 Aug 2014 17:40:05 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Sun, 24 Aug 2014 17:40:05 +0000 (UTC)
To: std-proposals@isocpp.org
Original-X-From: std-proposals+bncBCK2HM4L6YERB3OG5CPQKGQELW4VPCI@isocpp.org Sun Aug 24 19:40:00 2014
Return-path: <std-proposals+bncBCK2HM4L6YERB3OG5CPQKGQELW4VPCI@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-yk0-f200.google.com ([209.85.160.200])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBCK2HM4L6YERB3OG5CPQKGQELW4VPCI@isocpp.org>)
	id 1XLbln-0007ZE-6b
	for gclcip-std-proposals@m.gmane.org; Sun, 24 Aug 2014 19:39:59 +0200
Original-Received: by mail-yk0-f200.google.com with SMTP id 9sf41888545ykp.3
        for <gclcip-std-proposals@m.gmane.org>; Sun, 24 Aug 2014 10:39:58 -0700 (PDT)
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:mime-version:subject:from:in-reply-to:date
         :message-id:references:to:x-original-sender
         :x-original-authentication-results:reply-to:precedence:mailing-list
         :list-id:list-post:list-help:list-archive:list-subscribe
         :list-unsubscribe:content-type:content-transfer-encoding;
        bh=AGmDu/95slBgNVzAzPyzN9bbxMMLGOSVwRNR+6bZizA=;
        b=PPBgSv4EP89hDVfhr4/mgkyhfAZ5TPH864ziLIyR16933l5i3gt4MTy8jctNUUTT0b
         E58WJaca/PoP95tygI73ugPcllp2J3mcMjmNw2ckKcY3ZqfHjFobYZ0lHKQ3x0jU+JJY
         8bIesq2YEU4asnMjwIboSeQY6Ii271Kdymy+vf8kMOJrliONNFiJIUdTpfD/WSZwKidP
         emZzki/7xQ/zgJQ3DIuSiGxzWUK/7NVkmpugttfiKTLLIAAybvpjkaG/SqdvjXSQZjG2
         p1dxftVPUzsceRptxV+/NWfROcAcXAgmainFQt14VL+ZHNcLdfwqt9MpCVTDAze1uL/x
         1DOw==
X-Gm-Message-State: ALoCoQkAhrbhr5IbF0YEnYnCfM+9MwDSa/+xjVNnW8elm6gQbhfBR2BRvOeO2zG/Bh5wuu1Kfx+Z
X-Received: by 10.236.227.169 with SMTP id d39mr31045413yhq.45.1408901998106;
        Sun, 24 Aug 2014 10:39:58 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.50.234.161 with SMTP id uf1ls880260igc.30.gmail; Sun, 24 Aug
 2014 10:39:57 -0700 (PDT)
X-Received: by 10.50.134.106 with SMTP id pj10mr10769158igb.25.1408901997088;
        Sun, 24 Aug 2014 10:39:57 -0700 (PDT)
Original-Received: from mail-ig0-x232.google.com (mail-ig0-x232.google.com [2607:f8b0:4001:c05::232])
        by mx.google.com with ESMTPS id f16si1462312igz.53.2014.08.24.10.39.57
        for <std-proposals@isocpp.org>
        (version=TLSv1 cipher=ECDHE-RSA-RC4-SHA bits=128/128);
        Sun, 24 Aug 2014 10:39:57 -0700 (PDT)
Received-SPF: pass (google.com: domain of howard.hinnant@gmail.com designates 2607:f8b0:4001:c05::232 as permitted sender) client-ip=2607:f8b0:4001:c05::232;
Original-Received: by mail-ig0-f178.google.com with SMTP id uq10so1796807igb.17
        for <std-proposals@isocpp.org>; Sun, 24 Aug 2014 10:39:57 -0700 (PDT)
X-Received: by 10.42.62.6 with SMTP id w6mr20035196ich.24.1408901996912;
        Sun, 24 Aug 2014 10:39:56 -0700 (PDT)
Original-Received: from [10.0.1.2] (cpe-24-58-234-96.twcny.res.rr.com. [24.58.234.96])
        by mx.google.com with ESMTPSA id yt6sm26915815igb.10.2014.08.24.10.39.55
        for <std-proposals@isocpp.org>
        (version=TLSv1 cipher=ECDHE-RSA-RC4-SHA bits=128/128);
        Sun, 24 Aug 2014 10:39:55 -0700 (PDT)
In-Reply-To: <53F989E9.5090701@gmx.net>
X-Mailer: Apple Mail (2.1878.6)
X-Original-Sender: howard.hinnant@gmail.com
X-Original-Authentication-Results: mx.google.com;       spf=pass (google.com:
 domain of howard.hinnant@gmail.com designates 2607:f8b0:4001:c05::232 as
 permitted sender) smtp.mail=howard.hinnant@gmail.com;       dkim=pass
 header.i=@gmail.com;       dmarc=pass (p=NONE dis=NONE) header.from=gmail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>,
 <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:12464
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/12464>

I wrote:

>> class sha256
>> {
>>    SHA256_CTX state_;
>> public:=20
>>    static constexpr std::endian endian =3D std::endian::big;


On Aug 24, 2014, at 2:44 AM, Jens Maurer <Jens.Maurer@gmx.net> wrote:

> I agree there is a choice for std::uhash<> when
> converting a, say, "int" to a sequence of "unsigned char"s whether to
> perform endian conversion or not, and that influences the cross-platform
> reproducibility of the hash when e.g. hashing a sequence of "int"
> values.

This is the only role of the endian specifier, which can be any one of:

   static constexpr std::endian endian =3D std::endian::big;
   static constexpr std::endian endian =3D std::endian::little;
   static constexpr std::endian endian =3D std::endian::native;

std::endian::native will be equal to one of std::endian::little or std::end=
ian::big.  These enums are nothing more than the C++ version of the already=
 popular macros:  __BYTE_ORDER__, __ORDER_LITTLE_ENDIAN__, __ORDER_BIG_ENDI=
AN__.

And endian is really not even used directly by the hash functor uhash<>:

template <class Hasher =3D acme::siphash>
struct uhash
{
    using result_type =3D typename Hasher::result_type;

    template <class T>
    result_type
    operator()(T const& t) const noexcept
    {
        Hasher h;
        hash_append(h, t);
        return static_cast<result_type>(h);
    }
};

Instead there are now two traits:

template <class T> struct is_uniquely_represented;
template <class T, class HashAlgorithm> struct is_contiguously_hashable;

The first, is_uniquely_represented, is exactly what N3980 called is_contigu=
ously_hashable.  And is_contiguously_hashable has evolved into:

template <class T, class HashAlgorithm>
struct is_contiguously_hashable
    : public std::integral_constant<bool, is_uniquely_represented<T>{} &&
                                      (sizeof(T) =3D=3D 1 ||
                                       HashAlgorithm::endian =3D=3D endian:=
:native)>
{};

That is, whether or not a type T is contiguously hashable depends not only =
on if the type is_uniquely_represented, but also on whether the HashAlgorit=
hm needs to reverse the bytes of T before consuming it.  If the HashAlgorit=
hm's *requested endian* (HashAlgorithm::endian) is the same as the platform=
's *native endian* (endian::native), then the bytes do not need to be rever=
sed prior to feeding T to the HashAlgorithm.  And thus if T is also uniquel=
y represented, then one can feed T (or an array of T's) directly to the Has=
hAlgorithm.  This is the job of this function:

template <class Hasher, class T>
inline
std::enable_if_t
<
    is_contiguously_hashable<T, Hasher>{}
>
hash_append(Hasher& h, T const& t) noexcept
{
    h(std::addressof(t), sizeof(t));
}

If we are dealing with a platform/HashAlgorithm disagreement in endian, the=
n an alternative hash_append can be used for scalars:

template <class Hasher, class T>
inline
std::enable_if_t
<
    !is_contiguously_hashable<T, Hasher>{} &&
    (std::is_integral<T>{} || std::is_pointer<T>{} || std::is_enum<T>{})
>
hash_append(Hasher& h, T t) noexcept
{
    detail::reverse_bytes(t);
    h(std::addressof(t), sizeof(t));
}

Although it doesn't look like it in the source code, reverse_bytes is caref=
ully crafted to compile down to x86 instructions such as bswapl, at least o=
n clang/OSX, and so is maximally efficient.

So in summary, when uhash<HashAlgorithm> says:

        hash_append(h, t);

A correct hash_append will be chosen (at compile time), which for scalar ty=
pes that have endian issues, may or may not reverse the bytes of t dependin=
g on what the hash algorithm has requested, and what the platform's native =
endian is.

Most hash functions won't care about the endian of the scalars fed to them,=
 and they can indicate this by requesting the native endian, whatever that =
is:

   static constexpr std::endian endian =3D std::endian::native;

Given an underlying C implementation of a hash algorithm (such as SHA-256),=
 it would be quite easy to write adaptors around that C code with varying e=
ndian requirements, so that different parts of your code could use SHA-256 =
but reverse, or not reverse scalars as required:

class sha256
{
    SHA256_CTX state_;
public:=20
    static constexpr xstd::endian endian =3D xstd::endian::native;
    // ...


class sha256_little
{
    SHA256_CTX state_;
public:=20
    static constexpr xstd::endian endian =3D xstd::endian::little;
    // ...

// ...

uhash<sha256> h1;  // don't worry about endian
uhash<sha256_little> h2;  // ensure scalars are little endian prior to hash=
ing

Finally note that the implementation of hash_append is made simpler by the =
use of the (const void*, size_t) interface, as opposed to a (const unsigned=
 char*, size_t) interface.  With the latter, one would have to code:

template <class Hasher, class T>
inline
std::enable_if_t
<
    is_contiguously_hashable<T, Hasher>{}
>
hash_append(Hasher& h, T const& t) noexcept
{
    h(reinterpret_cast<const unsigned char*>(std::addressof(t)), sizeof(t))=
;
}

See https://github.com/HowardHinnant/hash_append for complete code.

Howard

--=20

---=20
You received this message because you are subscribed to the Google Groups "=
ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an e=
mail to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposa=
ls/.

.
