Path: csiph.com!eternal-september.org!reader02.eternal-september.org!.POSTED!not-for-mail
From: Tim Rentsch
Newsgroups: comp.lang.c++
Subject: Re: Can anyone improve this ?
Date: Mon, 17 Jan 2022 16:16:18 -0800
Organization: A noiseless patient Spider
Lines: 129
Message-ID: <86wnixrj8t.fsf@linuxsc.com>
References: <86ee56rzbz.fsf@linuxsc.com>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Injection-Info: reader02.eternal-september.org; posting-host="d40b94b342f96ad99b493372f9984bf0"; logging-data="5073"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX1/ExPZA50zxx1cvjj93ro/iIb5b4XaBfF8="
User-Agent: Gnus/5.11 (Gnus v5.11) Emacs/22.4 (gnu/linux)
Cancel-Lock: sha1:fxYX5URh/iC4tIKOg2ZGqHBviY4= sha1:oPr5XlEBqOCGY+oSw2GBZjre7V0=
Xref: csiph.com comp.lang.c++:82812
"Alf P. Steinbach" writes:
> On 17 Jan 2022 19:28, Tim Rentsch wrote:
>
>> Bonita Montero writes:
>>
>>> I had an idea how to write a fast UTF-8 strlen, here it is:
>>>
>>> size_t utf8Strlen( char const *str )
>>> {
>>> struct encode_t { size_t lenIncr, strIncr; };
>>> static encode_t const encodes[] =
>>> {
>>> { 1, 1 },
>>> { 0, 0 },
>>> { 1, 2 },
>>> { 1, 3 },
>>> { 1, 4 },
>>> { 0, 0 },
>>> { 0, 0 },
>>> { 0, 0 },
>>> { 0, 0 }
>>> };
>>> size_t len = 0;
>>> for( unsigned char c; (c = *str); )
>>> {
>>> encode_t const &enc =
>>> encodes[(size_t)countl_zero( ~c )];
>>> if( !enc.lenIncr ) [[unlikely]]
>>> return -1;
>>> len += enc.lenIncr;
>>> for( char const *cpEnd = str + enc.strIncr; ++str != cpEnd; )
>>> if( ((unsigned char)*str & 0x0C0) != 0x080 ) [[unlikely]]
>>> return -1;
>>> }
>>> return len;
>>> }
>>>
>>> Has anyone further ideas to improve this ?
>>
>> Just for fun -
>>
>> size_t
>> utf8_units_two( const char *s ){
>> size_t r = -1;
>> unsigned char c;
>>
>> next:
>> switch( r++, c = *s++, c >> 3 ){
>> cases(16,17,18,19,20,21,22,23,31): return -1;
>>
>> cases(30): if( c = *s++, c >> 6 != 2 ) return -1;
>> cases(28,29): if( c = *s++, c >> 6 != 2 ) return -1;
>> cases(24,25,26,27): if( c = *s++, c >> 6 != 2 ) return -1;
>> cases(1,2,3,4,5,6,7,8,9,10,11,12,13,14,15): goto next;
>>
>> cases(0): if( c != 0 ) goto next;
>> }
>>
>> return r;
>> }
>
> That's very ugly.
Beauty is in the eye of the beholder. Personally I think it's
rather pretty.
>> (Note: the cases() macro produces one 'case X:' for each
>> argument X except the last one where the ':' is omitted,
>> written using the standard C preprocessor, and left as an
>> exercise for any ambitious readers.)
>
> Exercise for you: test that unrevealed macro with Visual C++.
I don't have access to a Visual C++ system. The macro I wrote
compiles under gcc and g++ as C99, C11, C++11, C++14, and C++17,
using a -pedantic-errors option. (Using clang and clang++ gives
the same results.)
> (Making it work also with Visual C++ is doable.)
Does this need more than just writing the definition in
conforming C11 or C++11?
>> Incidentally, this code provides the only example I remember
>> where using 'goto' is pretty much unavoidable, in that any
>> re-writing without 'goto' seems awkward or inferior in some
>> other way. I guess I should say, at least not that I could
>> find, maybe someone else can do better.
>
> The `for(;;)`, `continue;` and `break;` constructs come to mind, so
> you're right, someone else can do better.
It is of course possible to write the function without using any
'goto's. Whether such a writing is an improvement is another
matter. To my eye the goto version expresses the structure more
directly and in a form that is easier to comprehend. I did try
actually writing a version with a continue-able loop around the
switch(), but it looked clumsy and not as easy to follow as the
goto version. If you have an example to post I definitely would
like to see it.
> At a guess the code counts the number of Unicode code points?
Yes, sorry, I thought that was apparent from the earlier
context.
> Does it return -1 also for "too long" UTF-8 sequence?
The -1 return values is for character sequences that are not
syntactically well-formed for UTF-8.
>> This code also has the interesting property that along the main
>> code path there are no conditional branches (as compiled by gcc
>> under -O2).
>
> Smart compiler removed all the `if`'s?
>
> I guess you're flame-baiting a little just for fun, but I chose to
> take it seriously.
I'm not flame-baiting at all. By "main code path" I mean that
part of the code that deals with normal ASCII characters (and not
counting the terminating null character). The switch() statement
generates an unconditional indirect branch, indexing a table of
32 targets, and the 15 targets for normal ASCII characters simply
go around the loop again, with no tests. I expect you can see
that based on the line with 15 cases, which just does a 'goto'
to start the loop again.