Follow @Openwall on Twitter for new release announcements and other news
[<prev] [next>] [<thread-prev] [thread-next>] [day] [month] [year] [list]
Message-ID: <20261008140056.GM23438@brightrain.aerifal.cx>
Date: Thu, 8 Oct 2026 10:00:56 -0400
From: Rich Felker <dalias@...c.org>
To: musl@...ts.openwall.com
Subject: Re: fnmatch: invalid multibyte sequence in the string is
 taken as end of string ("" matches "\xff" in UTF-8 locales)

On Thu, Oct 08, 2026 at 02:27:31PM +1100, raf wrote:
> On Tue, Oct 06, 2026 at 04:01:53AM +0000, "He, Zhibo" <zhibo.he25@...erial.ac.uk> wrote:
> 
> > Hello,
> > 
> > On musl 1.2.6 in a UTF-8 locale, fnmatch() returns 0 whenever the pattern
> > is exhausted and the next bytes of the string form an invalid multibyte
> > sequence. fnmatch("", "\xff", 0) returns 0, and so does
> > fnmatch("admin", "admin\xff/../../etc/passwd", 0). glibc returns
> > FNM_NOMATCH for both, and so does musl in the C locale.
> > 
> > In src/regex/fnmatch.c, str_next() returns 0 at end of string and -1 for an
> > illegal sequence. fnmatch_internal() tests k <= 0 after two of its
> > str_next() calls, in the head loop and in the tail check, so -1 takes the
> > end-of-string branch and the function returns 0 when the pattern is also at
> > its end. With a star in the pattern only trailing bytes 0x80-0xBF reach the
> > tail check: fnmatch("*.txt", "x.txt\x80\x80", 0) returns 0 and
> > fnmatch("*.txt", "x.txt\xff", 0) does not. Both tests are unchanged on
> > master (9b2d8a16).
> > 
> > To reproduce:
> > 
> > #include <fnmatch.h>
> > #include <locale.h>
> > #include <stdio.h>
> > int main(void)
> > {
> >       setlocale(LC_ALL, "");                   /* C.UTF-8 on Alpine, LANG unset */
> >       printf("%d\n", fnmatch("", "\xff", 0));               /* 0 = match */
> >       printf("%d\n", fnmatch("admin", "admin\xff", 0));     /* 0 = match */
> >       printf("%d\n", fnmatch("admin", "admin\xc3\xa9", 0)); /* valid e-acute: 1 */
> >       printf("%d\n", fnmatch("*.txt", "x.txt\x80\x80", 0)); /* 0 = match */
> >       return 0;
> > }
> > 
> > On Alpine (musl 1.2.6-r2) with LANG unset this prints 0, 0, 1, 0. Under
> > LC_ALL=C it prints 1, 1, 1, 1, and so does glibc 2.39 under C.UTF-8. On
> > Alpine the over-match shows through glob(), busybox ash (case and pathname
> > expansion) and PHP's fnmatch().
> > 
> > raf's patch of 4 October 2025, "fnmatch: Make ? match binary/non-character
> > byte (like * does)", changes the first of the two tests so that "?" accepts
> > an invalid byte. It keeps the END case, and with it applied the program
> > above still prints 0, 0, 1, 0.
> > 
> > The patch below, against master, treats -1 as a mismatch at both sites.
> > With the patched file linked against musl 1.2.6-r2 the program prints
> > 1, 1, 1, 1, and fnmatch("*", "\xff", 0) still returns 0. Its first hunk
> > changes the line that raf's patch changes. If "?" should match an invalid
> > byte, the exception belongs in that k < 0 branch.
> > 
> > --- a/src/regex/fnmatch.c
> > +++ b/src/regex/fnmatch.c
> > @@ -181,7 +181,9 @@
> >                   break;
> >             default:
> >                   k = str_next(str, n, &sinc);
> > -                 if (k <= 0)
> > +                 if (k < 0)
> > +                       return FNM_NOMATCH;
> > +                 if (!k)
> >                         return (c==END) ? 0 : FNM_NOMATCH;
> >                   str += sinc;
> >                   n -= sinc;
> > @@ -242,7 +244,7 @@
> >             c = pat_next(p, endpat-p, &pinc, flags);
> >             p += pinc;
> >             if ((k = str_next(s, endstr-s, &sinc)) <= 0) {
> > -                 if (c != END) return FNM_NOMATCH;
> > +                 if (c != END || k < 0) return FNM_NOMATCH;
> >                   break;
> >             }
> >             s += sinc;
> > 
> > Found by differential testing of musl, glibc and OpenBSD in an MSc project
> > at Imperial College London.
> > 
> > Zhibo He
> 
> Hi Zhibo,
> 
> That's interesting. I've made more changes to my local copy of musl's fnmatch()
> so that it supports non-UTF8 bytes (not all filenames are UTF-8, and they're not
> necessarily all even text). With my current version, your test program prints
> 1, 1, 1, 1. Its str_next() never returns -1. A patch is attached.

Patterns are specified to match characters, not arbitrary bytes.
Putting byte strings that are not valid characters into filenames is
generally not a supported usage (by musl or by the standards).

The choice to have * match spans that potentially contain invalid
bytes was made a long time ago based on several considerations:

- Not spending time decoding every character to ensure it's valid when
  checking the match.

- Because * doesn't match a specific number of characters, there were
  not really any weird corner cases that might have wrong behavior
  from ambiguity of what unit of input should count for a match. OTOH
  having ? or brackets match bogus bytes/partial characters has weird
  consequences like s1 matching p1 and s2 matching p2 but s1s2 not
  matching p1p2.

- Avoiding gratuitous difficulty cleaning up after accidentally
  extracting an archive with bad encoding that wasn't interpreted
  properly during extraction.

- Originally at the time it was written, not having a byte-based C
  locale that would allow doing these things just by temporarily
  setting LC_CTYPE=C.

Trying to be clever with invalid byte sequences in UTF-8 rather than
treating them as invalid is generally a recipe for bad things
happening.

Powered by blists - more mailing lists

Confused about mailing lists and their use? Read about mailing lists on Wikipedia and check out these guidelines on proper formatting of your messages.