Follow @Openwall on Twitter for new release announcements and other news
[<prev] [<thread-prev] [day] [month] [year] [list]
Message-ID: <20261009130824.GP23438@brightrain.aerifal.cx>
Date: Fri, 9 Oct 2026 09:08:25 -0400
From: Rich Felker <dalias@...c.org>
To: musl@...ts.openwall.com
Subject: Re: fnmatch: invalid multibyte sequence in the string is
 taken as end of string ("" matches "\xff" in UTF-8 locales)

On Thu, Oct 08, 2026 at 02:27:31PM +1100, raf wrote:
> That's interesting. I've made more changes to my local copy of musl's fnmatch()
> so that it supports non-UTF8 bytes (not all filenames are UTF-8, and they're not
> necessarily all even text). With my current version, your test program prints
> 1, 1, 1, 1. Its str_next() never returns -1. A patch is attached.

Getting back to the technical aspects of this, I think you should
check how your patch makes ${foo%?} behave. I would expect shells
implement this functionality by stepping back from the end of the
string one byte at a time and calling fnmatch until they get a
successful match. If this is the case, your patch will make them rip
apart a final multibyte character and produce malformed output.

Note that this does not happen with *. The shortest match of a lone *
is always the empty string, so ${foo%*} is just $foo. And the longest
match of a lone * is the full string. And if you have anything other
than a * prior to the *, it will necessarily anchor to a character
boundary -- as long as patterns matching a single character only match
a single character and don't wrongly match a single byte of a
multibyte character.

The same should also apply from the beginning of a string using
${foo#?} etc.

Rich

Powered by blists - more mailing lists

Confused about mailing lists and their use? Read about mailing lists on Wikipedia and check out these guidelines on proper formatting of your messages.