220 12556 <53FE7CFB.1090509@gmail.com> article
Path: news.gmane.org!not-for-mail
From: Miro Knejp <miro.knejp@gmail.com>
Newsgroups: gmane.comp.lang.c++.isocpp.proposals
Subject: Re: Cryptographic hash functions reloaded [was
 Interest in cryptographic functions within the standard library]
Date: Thu, 28 Aug 2014 02:51:07 +0200
Lines: 164
Approved: news@gmane.org
Message-ID: <53FE7CFB.1090509@gmail.com>
References: <53F8F620.3090701@gmx.de> <B33BCE97-51A4-4880-9C26-5347D6C8A02A@gmail.com> <53F989E9.5090701@gmx.net> <B0F486D7-A068-4A30-8D5E-69E826F1C943@gmail.com> <53FE5E8A.2090003@gmx.net>
Reply-To: std-proposals@isocpp.org
NNTP-Posting-Host: plane.gmane.org
Mime-Version: 1.0
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
X-Trace: ger.gmane.org 1409187084 15077 80.91.229.3 (28 Aug 2014 00:51:24 GMT)
X-Complaints-To: usenet@ger.gmane.org
NNTP-Posting-Date: Thu, 28 Aug 2014 00:51:24 +0000 (UTC)
To: std-proposals@isocpp.org
Original-X-From: std-proposals+bncBC6ONSXJ54LBBA727GPQKGQEB7OOCBI@isocpp.org Thu Aug 28 02:51:16 2014
Return-path: <std-proposals+bncBC6ONSXJ54LBBA727GPQKGQEB7OOCBI@isocpp.org>
Envelope-to: gclcip-std-proposals@m.gmane.org
Original-Received: from mail-la0-f69.google.com ([209.85.215.69])
	by plane.gmane.org with esmtp (Exim 4.69)
	(envelope-from <std-proposals+bncBC6ONSXJ54LBBA727GPQKGQEB7OOCBI@isocpp.org>)
	id 1XMnvo-0002ls-2Z
	for gclcip-std-proposals@m.gmane.org; Thu, 28 Aug 2014 02:51:16 +0200
Original-Received: by mail-la0-f69.google.com with SMTP id b8sf367530lan.0
        for <gclcip-std-proposals@m.gmane.org>; Wed, 27 Aug 2014 17:51:15 -0700 (PDT)
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20130820;
        h=x-gm-message-state:message-id:date:from:user-agent:mime-version:to
         :subject:references:in-reply-to:x-original-sender
         :x-original-authentication-results:reply-to:precedence:mailing-list
         :list-id:list-post:list-help:list-archive:list-subscribe
         :list-unsubscribe:content-type;
        bh=QD0kN8bcyR0E5OnZ4c0coGQudx7eIe4/AHGIzjDFuqg=;
        b=gwCqIVHdB/qtP46VG0oUFNieOk+DboU8bJCnUm/vssznlDyZhZa1He/Ei9S+E/yuSx
         eyXLlkG8fyXPYLF9dmOBdCAu0jjTvZVrp9clNs4Z95fzJLlLKPUCKDtXEwMysr4J1aNk
         cmvXi8mcvyT+DXjnVkR2dkFwOU2mlfKvYsddcu2C/Be50yKuUgv4wAErbzvCUrKxe1B5
         GmHnMRBi5Y3uCioS2PLk7HbqF3g/HB3606QZ1v5ZH+E/Q2wFBL5w2XqMUpLUYtxXL59Z
         P5Pre0hI4G1OC/nE4ZXAiZjpZUZb50yNeFrU4IzaeQa17bfy/+pJ3Pbp1sGsCFHsM8e5
         0amQ==
X-Gm-Message-State: ALoCoQnMDtQ3fTQaswJJmp5MpXBe7q2G+m5LXNBsm7R6/OhvlimY/1eTtDR9XvKTi9Qv+mw3Sh0f
X-Received: by 10.180.36.98 with SMTP id p2mr78073wij.0.1409187075707;
        Wed, 27 Aug 2014 17:51:15 -0700 (PDT)
X-BeenThere: std-proposals@isocpp.org
Original-Received: by 10.180.206.227 with SMTP id lr3ls94308wic.18.gmail; Wed, 27 Aug
 2014 17:51:14 -0700 (PDT)
X-Received: by 10.180.92.73 with SMTP id ck9mr33706704wib.54.1409187074616;
        Wed, 27 Aug 2014 17:51:14 -0700 (PDT)
Original-Received: from mail-wi0-x229.google.com (mail-wi0-x229.google.com [2a00:1450:400c:c05::229])
        by mx.google.com with ESMTPS id k10si4105573wjx.87.2014.08.27.17.51.14
        for <std-proposals@isocpp.org>
        (version=TLSv1 cipher=ECDHE-RSA-RC4-SHA bits=128/128);
        Wed, 27 Aug 2014 17:51:14 -0700 (PDT)
Received-SPF: pass (google.com: domain of miro.knejp@gmail.com designates 2a00:1450:400c:c05::229 as permitted sender) client-ip=2a00:1450:400c:c05::229;
Original-Received: by mail-wi0-f169.google.com with SMTP id n3so133wiv.0
        for <std-proposals@isocpp.org>; Wed, 27 Aug 2014 17:51:14 -0700 (PDT)
X-Received: by 10.180.84.66 with SMTP id w2mr32465581wiy.27.1409187074352;
        Wed, 27 Aug 2014 17:51:14 -0700 (PDT)
Original-Received: from [192.168.42.33] (ppp-93-104-164-220.dynamic.mnet-online.de. [93.104.164.220])
        by mx.google.com with ESMTPSA id cy10sm5338105wjb.21.2014.08.27.17.51.12
        for <std-proposals@isocpp.org>
        (version=TLSv1 cipher=ECDHE-RSA-RC4-SHA bits=128/128);
        Wed, 27 Aug 2014 17:51:13 -0700 (PDT)
User-Agent: Mozilla/5.0 (Windows NT 6.3; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.6.0
In-Reply-To: <53FE5E8A.2090003@gmx.net>
X-Original-Sender: miro.knejp@gmail.com
X-Original-Authentication-Results: mx.google.com;       spf=pass (google.com:
 domain of miro.knejp@gmail.com designates 2a00:1450:400c:c05::229 as
 permitted sender) smtp.mail=miro.knejp@gmail.com;       dkim=pass
 header.i=@gmail.com;       dmarc=pass (p=NONE dis=NONE) header.from=gmail.com
Precedence: list
Mailing-list: list std-proposals@isocpp.org; contact std-proposals+owners@isocpp.org
List-ID: <std-proposals.isocpp.org>
X-Google-Group-Id: 399137483710
List-Post: <http://groups.google.com/a/isocpp.org/group/std-proposals/post>, <mailto:std-proposals@isocpp.org>
List-Help: <http://support.google.com/a/isocpp.org/bin/topic.py?topic=25838>, <mailto:std-proposals+help@isocpp.org>
List-Archive: <http://groups.google.com/a/isocpp.org/group/std-proposals/>
List-Subscribe: <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>,
 <mailto:std-proposals+subscribe@isocpp.org>
List-Unsubscribe: <mailto:googlegroups-manage+399137483710+unsubscribe@googlegroups.com>,
 <http://groups.google.com/a/isocpp.org/group/std-proposals/subscribe>
Xref: news.gmane.org gmane.comp.lang.c++.isocpp.proposals:12556
Archived-At: <http://permalink.gmane.org/gmane.comp.lang.c++.isocpp.proposals/12556>

Maybe this whole issue simply shows that algorithms with strict size 
requirements should not be defined in terms of char, short, int but 
int8_t, int16_t, and so on. If hash_append has (by default) only 
overloads for exactly sized types then the compiler should pick the 
correct one when the user feeds it unsized types like int. On machines 
where no int8_t exists no hash_append overload for int8_t exists.

Where the same source code is to produce equal hashes for the same data 
structures on different machines then the only portable way is to use 
the explicitly sized types. That is honestly the *only* way to be sure 
of consistent hash values and should maybe be added as a note somewhere.

Though I think "char" (not "signed char" or "unsigned char" as those are 
3 different types) should be treated implementation-defined as the 
standard uses this type only for character values in strings. On a 
machine where char has more than 8 bit only the implementation knows 
which of these bits are representative of the value of a string 
character/codepoint. Same goes for wchar_t.

Regarding the void* question, I think Howard's is_contiguously_hashable 
covers this nicely. If your type has no padding in it, specialize the 
trait. If it does, provide your own hash_append overload and feed it the 
required members. I think a good set of predefined overloads for 
hash_append would be

hash_append(Hasher&, const T&) // enable if is_contiguously_hashable<T, 
Hasher> is true
hash_append(Hasher&, array_view<T>) // enable if hash_append(Hasher, T) 
is well-formed
hash_append(Hasher&, basic_string_view<Char, Traits>) // enable if 
hash_append(Hasher, Char) is well-formed

Having is_contiguously_hashable<T, Hasher> predefined as true for the 
types [u]int[8|16|32|64]_t, char, wchar_t and char[16|32]_t allows the 
implementation to select which integral types are acceptable and can 
then internally provide specializations with tag dispatching. Using 
unsized types like short or int would pick the proper overload depending 
on what int##_t typedefs alias to.

A further alternative would be to provide a third parameter in the form 
hash_append(Hasher&, const T&, valid_bits<N>), so even on architectures 
without direct int8_t support one could use ints for storage and only 
mask the leading N bits as relevant for the hash.

This should make any need for a hash_append(void*, size_t) overload 
obsolete. The Hasher itself needs a (void* p, size_t n) overload where n 
denotes the number of valid OCTETS pointed to by p. hash_append() has 
then already taken care of endianess and other details by applying 
Hasher's traits. The nice hing here is that if a type has no padding and 
the user *decided that endianess, etc. does not matter* then enabling 
is_contiguously_hashable makes hash_append() feed the entire structure 
to (void*, size_t). Appropriate warning signs should be positioned 
around is_contiguously_hashable to make the user aware of its positive 
and negative implications.

Does this make sense?

Am 28.08.2014 00:41, schrieb Jens Maurer:
> Hi Howard!
>
> I'll have to ask a few more questions here.  If something gets
> standardized in this area, I'd like to see a roadmap how all
> platforms supported by C++ can get portable (crypto) hash values,
> even if not maximally efficient on "strange" environments.
>
> I'm assuming that (in an abstract sense) a crypto hash algorithm
> (and most others) hash a sequence of octets (i.e. 8-bit quantities).
>
> (SHA256 and others can do odd trailing bits, but let's ignore
> this for now.)
>
> I presume I'm passing these octets to the hash algorithm using an
> array of (unsigned) char, right?
>
> Is this assumption also true for platforms where (unsigned) char
> is e.g. 32-bits (DSPs)?  If so, the following implementation doesn't
> work there, because it makes no effort to split e.g. T==int (suppose
> it's 32-bit) into four individual (unsigned) char objects.
>
> (Oh, and on such platforms, 1 == sizeof(char) == sizeof(int).)
>
>> hash_append(Hasher& h, T const& t) noexcept
>> {
>>      h(std::addressof(t), sizeof(t));
>> }
>
> On 08/24/2014 07:39 PM, Howard Hinnant wrote:
>> If we are dealing with a platform/HashAlgorithm disagreement in endian, then an alternative hash_append can be used for scalars:
> So, the endianness is a boolean, not a three-way type?  Either you're "native" or not
> seems all that matters, from the code you presented.
>
>
> I think we're mixing two slightly related, but distinct aspects here.
>
> One aspect is the fact that hashing an array of "short" values with a portable
> hash (such as sha256) should give the same value across all platforms.
> (The cost of endianess normalization is negligible compared to the cost
> of running sha256 at all.)  Maybe we need to identify those hash algorithms
> that are supposed to give portable results.
>
> So, for example, hashing this:
>
>    short a[] = { 1, 2, 3, 4 };
>
> with sha256 should give consistent cross-platform results.
> Thus, when feeding octets to the hash algorithm, we must have a
> (default) idea how to map a "short" value into (two? four? more?)
> octets.  That might be big or little endian, or something entirely
> different (e.g. a variable-length encoding, which might actually
> turn out to be the most portable).
>
>
> The other aspect is the fact that hash algorithms such as sha256 like
> to process e.g. 32-bits (= 4 octets) at once.  When reading four
> octets from memory, it's helpful to be able to simply read them
> into a register on "suitable" platforms and only do the endianess
> (or other) conversion on the remainder of the platforms.  But, on
> the abstract level, this is not a configuration option, it's a
> question of correctness.
>
> [Giving the user a way to opt-out of this endianess correctness is
> fine for me (emphasis: opt-out).]
>
>
> I don't think a single "endian" value captures both aspects.
>
>
>> uhash<sha256> h1;  // don't worry about endian
>> uhash<sha256_little> h2;  // ensure scalars are little endian prior to hashing
> The simple name must result in the portable hash value.
>
>> Finally note that the implementation of hash_append is made simpler by the use of the (const void*, size_t) interface, as opposed to a (const unsigned char*, size_t) interface.  With the latter, one would have to code:
>>
>> template <class Hasher, class T>
>> inline
>> std::enable_if_t
>> <
>>      is_contiguously_hashable<T, Hasher>{}
>> hash_append(Hasher& h, T const& t) noexcept
>> {
>>      h(reinterpret_cast<const unsigned char*>(std::addressof(t)), sizeof(t));
>> }
> I agree "void *" is simpler, but I continue to believe this is a more dangerous
> interface, allowing to inadvertently pass stuff with padding in it.  Note
> that the user is not expected to write this code, but rely on the standard
> library to hash scalar types.
>
> (If hashing comes up in Urbana-Champaign, please grab me so that I can voice
> a "strongly against" for this particular aspect.)
>
> (You can use two static_casts via "void *" if the reinterpret_cast is too
> dreadful.)
>
> Jens
>

-- 

--- 
You received this message because you are subscribed to the Google Groups "ISO C++ Standard - Future Proposals" group.
To unsubscribe from this group and stop receiving emails from it, send an email to std-proposals+unsubscribe@isocpp.org.
To post to this group, send email to std-proposals@isocpp.org.
Visit this group at http://groups.google.com/a/isocpp.org/group/std-proposals/.

.
