Follow @Openwall on Twitter for new release announcements and other news
[<prev] [next>] [thread-next>] [day] [month] [year] [list]
Message-ID: <20260903162221.GA24252@brightrain.aerifal.cx>
Date: Thu, 3 Sep 2026 12:22:22 -0400
From: Rich Felker <dalias@...c.org>
To: musl@...ts.openwall.com
Subject: Collation implementation is functional

Collation is working! This is the culmination of a long R&D project
that stretched out a lot longer than I would have preferred, but the
results are pretty much just what I wanted. Getting to the point of
having real functional code based on the designs so far is one of the
last big steps of the locale support overhaul project, funded by NLnet
and the NGI Zero Core Fund, and brings us to a good place for
switching over to the new locale system.


Locale utilities side:

The fractional-UCA root data parser is now wired up to the localedef
utility with a custom -C option to specify a file to read from, and it
now produces binary output for processing at runtime as part of the
memory-mapped locale.

There are a lot of enhancements this could still undergo. It does not
support increasing or decreasing the number of weight levels (hard
coded at 3), uses a built-in dictionary of common patterns rather than
building one from a histogram of the loaded data, and does not yet
have any functionality to transform the root data applying
locale-specific tailorings.

The problem of tailoring could be solved either by directly parsing
and applying the tailorings, or with a preprocessing step using
existing or new tooling. I'm probably not going to be pursuing this
right away, as it's not on the dependency path for testing and
integrating collation support in musl, but it needs to happen at some
point as part of making locales for musl to use.

Source is at https://codeberg.org/dalias/musl-locale-tools-draft and
builds from the included makefile.



Runtime side:

Both basic proof of concept (iteration of collation elements from an
input string) and a working strxfrm (transformation of string to a
null-terminated binary sort key) are working, and have been
smoke-tested manually for basic coverage:

- Straightforward one-to-one mappings
- Matching prefix contexts (middle dot following L)
- Non-matching prefix contexts (middle dot not following L)
- One-to-many mappings
- Two-to-one mappings
- Reordering mappings (Thai vowel)
- Auto-generated ideographic radical-stroke-order weights
- Implicit weights
- Reversed-order secondary weights

This list will make the basis for an automated smoke test. There are
also Unicode-provided test vectors which have not yet been tested, but
these do not assert specific behavior generating keys; they just
assert the correct ordering of a long list of short string fragments.

Under the hood, two layers of collation iterators have been
implemented and are functional. The first processes the sequence of
NFD codepoints from the source string and chooses the suitable
context-dependent mapping consuming one or more characters and
producing one or more collation elements. The second walks thru a
potentially multi-element sequence obtained from a mapping for
processing one at a time.

The direct comparison function strcoll has been written but not hooked
up or tested.

Source is at https://codeberg.org/dalias/musl-uca-draft and includes a
demo. UnicodeData.txt is needed to build it, and a locale file with
collation data built from the above localedef utility is needed to run
it.

Powered by blists - more mailing lists

Confused about mailing lists and their use? Read about mailing lists on Wikipedia and check out these guidelines on proper formatting of your messages.