Follow @Openwall on Twitter for new release announcements and other news
[<prev] [next>] [<thread-prev] [thread-next>] [day] [month] [year] [list]
Message-ID: <20261006161457.GL23438@brightrain.aerifal.cx>
Date: Tue, 6 Oct 2026 12:14:57 -0400
From: Rich Felker <dalias@...c.org>
To: "He, Zhibo" <zhibo.he25@...erial.ac.uk>
Cc: "musl@...ts.openwall.com" <musl@...ts.openwall.com>
Subject: Re: fnmatch: invalid multibyte sequence in the string is
 taken as end of string ("" matches "\xff" in UTF-8 locales)

On Tue, Oct 06, 2026 at 04:01:53AM +0000, He, Zhibo wrote:
> Hello,
> 
> On musl 1.2.6 in a UTF-8 locale, fnmatch() returns 0 whenever the pattern
> is exhausted and the next bytes of the string form an invalid multibyte
> sequence. fnmatch("", "\xff", 0) returns 0, and so does
> fnmatch("admin", "admin\xff/../../etc/passwd", 0). glibc returns
> FNM_NOMATCH for both, and so does musl in the C locale.
> 
> In src/regex/fnmatch.c, str_next() returns 0 at end of string and -1 for an
> illegal sequence. fnmatch_internal() tests k <= 0 after two of its
> str_next() calls, in the head loop and in the tail check, so -1 takes the
> end-of-string branch and the function returns 0 when the pattern is also at
> its end. With a star in the pattern only trailing bytes 0x80-0xBF reach the
> tail check: fnmatch("*.txt", "x.txt\x80\x80", 0) returns 0 and
> fnmatch("*.txt", "x.txt\xff", 0) does not. Both tests are unchanged on
> master (9b2d8a16).
> 
> To reproduce:
> 
> #include <fnmatch.h>
> #include <locale.h>
> #include <stdio.h>
> int main(void)
> {
>       setlocale(LC_ALL, "");                   /* C.UTF-8 on Alpine, LANG unset */
>       printf("%d\n", fnmatch("", "\xff", 0));               /* 0 = match */
>       printf("%d\n", fnmatch("admin", "admin\xff", 0));     /* 0 = match */
>       printf("%d\n", fnmatch("admin", "admin\xc3\xa9", 0)); /* valid e-acute: 1 */
>       printf("%d\n", fnmatch("*.txt", "x.txt\x80\x80", 0)); /* 0 = match */
>       return 0;
> }
> 
> On Alpine (musl 1.2.6-r2) with LANG unset this prints 0, 0, 1, 0. Under
> LC_ALL=C it prints 1, 1, 1, 1, and so does glibc 2.39 under C.UTF-8. On
> Alpine the over-match shows through glob(), busybox ash (case and pathname
> expansion) and PHP's fnmatch().
> 
> raf's patch of 4 October 2025, "fnmatch: Make ? match binary/non-character
> byte (like * does)", changes the first of the two tests so that "?" accepts
> an invalid byte. It keeps the END case, and with it applied the program
> above still prints 0, 0, 1, 0.
> 
> The patch below, against master, treats -1 as a mismatch at both sites.
> With the patched file linked against musl 1.2.6-r2 the program prints
> 1, 1, 1, 1, and fnmatch("*", "\xff", 0) still returns 0. Its first hunk
> changes the line that raf's patch changes. If "?" should match an invalid
> byte, the exception belongs in that k < 0 branch.

>From what I recall, the original intent was that only * be able to
match invalid encodings, mainly for the sake of being able to do
things like 'rm *'. This is supported by the logic for advancing on
failure in the comment near the bottom of fnmatch_internal. Skipping
invalid bytes here would break the ability to match them with '?', if
we wanted '?' to match them.

> --- a/src/regex/fnmatch.c
> +++ b/src/regex/fnmatch.c
> @@ -181,7 +181,9 @@
>                   break;
>             default:
>                   k = str_next(str, n, &sinc);
> -                 if (k <= 0)
> +                 if (k < 0)
> +                       return FNM_NOMATCH;
> +                 if (!k)
>                         return (c==END) ? 0 : FNM_NOMATCH;
>                   str += sinc;
>                   n -= sinc;
> @@ -242,7 +244,7 @@
>             c = pat_next(p, endpat-p, &pinc, flags);
>             p += pinc;
>             if ((k = str_next(s, endstr-s, &sinc)) <= 0) {
> -                 if (c != END) return FNM_NOMATCH;
> +                 if (c != END || k < 0) return FNM_NOMATCH;
>                   break;
>             }
>             s += sinc;
> 
> Found by differential testing of musl, glibc and OpenBSD in an MSc project
> at Imperial College London.

Thanks! I'll probably apply a version where the first hunk looks more
like the second, keeping the FNM_NOMATCH cases together, but otherwise
this looks good.

Rich

Powered by blists - more mailing lists

Confused about mailing lists and their use? Read about mailing lists on Wikipedia and check out these guidelines on proper formatting of your messages.