Follow @Openwall on Twitter for new release announcements and other news
[<prev] [next>] [<thread-prev] [thread-next>] [day] [month] [year] [list]
Message-ID: <20261004161557.GB3542221@port70.net>
Date: Sun, 4 Oct 2026 18:15:57 +0200
From: Szabolcs Nagy <nsz@...t70.net>
To: Rich Felker <dalias@...c.org>
Cc: Alex Rønne Petersen <alex@...xrp.com>,
	musl@...ts.openwall.com
Subject: Re: [PATCH] riscv: declare all vector registers as clobbers
 of syscalls

* Rich Felker <dalias@...c.org> [2026-10-04 11:08:15 -0400]:
> On Sun, Oct 04, 2026 at 02:21:35PM +0200, Szabolcs Nagy wrote:
> > * Szabolcs Nagy <nsz@...t70.net> [2026-10-04 12:55:13 +0200]:
> > > * Alex Rønne Petersen <alex@...xrp.com> [2026-10-04 04:48:52 +0200]:
> > > > The kernel intentionally clobbers vector registers (by setting them to all 1s).
> > > > musl didn't declare this, so with a compiler targeting the V extension, this
> > > > could lead to all sorts of breakage that at first glance looks like
> > > > miscompilations.
> > > 
> > > for the record the aarch64 behaviour is keeping the
> > > normal parts of the simd regs and zeroing the sve
> > > extension bits.
> > > 
> > > https://www.kernel.org/doc/html/latest/arch/arm64/sve.html#system-call-behaviour
> > > 
> > > so current musl clobber is wrong when compiled with +sve
> > > (note: "cc" is actually preserved not clobbered)
> > > 
> > > test to demonstrate:
> > > https://godbolt.org/z/xKfozYzjj
> > 
> > i dont see a clean solution to the d8..d15 problem:
> > these are call preserved fp regs so a clobber on the
> > overlapping sve z8..z15 spills them, the kernel clobbers
> > the top bits only but i dont see asm syntax for that.
> > 
> > either we explicitly build musl as +nosve or make
> > the syscalls noinline when building with sve
> > (normal call convention clobbers top bits of sve regs)
> 
> Is there any way musl could plausibly be using upper range of sve
> registers? This is probably at present only a theoretical problem, but

i compiled musl with armv8-a+sve and e.g.

	} else if (v[0]==AT_HWCAP3 || v[0]==AT_HWCAP4) {
		v[0] = AT_IGNORE;
		v[1] = 0;
	}

code from aarch64  __set_thread_area is compiled to

  index z31.d, #1, #-1
  ...
  str   q31, [x1]

the index fills z31 with the 64bit seq 1,0,-1,-2,...
then q31 stores the bottom 128bit of it.

since sve is richer isa than the old advsimd, the
compiler may use sve instructions even if it only
cares about the bottom 128bits because there is
no equivalent in the old isa.

there are not many sve intructions though and they
are mostly in math code where gcc prefers sve isa
to moving between fp and int regs e.g.

  orr   z31.d, z31.d, #0x7fe0000000000000

sets double prec exponents in all lanes of z31, but
the code is scalar so only uses the bottom lane.

if the question is "does musl benefit from sve" then
it's unlikely (some string functions may, but requires
hand optimized sve variant).

> I think we should fix it. Forcing nosve seems like the clean solution,
> but I'm not sure that helps in a theoretical case where someone is
> using LTO and the compiler tries to inline code from inside musl
> across other functions in musl (nosve) into a calling application
> (using sve). The only solution here short of fixing the badly designed
> compiler asm constraints might be forcing noinline if built to support
> LTO.

lto should take care of that (in practice there may
be bugs) i think noinline attribute can be broken too.

unfortunately there is no clean way to add nosve. gcc has

  #pragma GCC target "+nosve"

the ai said clang instead needs

  #pragma clang attribute push (__attribute__((target("no-sve"))), apply_to=function)

so either we use different compiler specific magic or
use cflags and override user -march settings via
-march=armv8-a or -march=armv9-a+nosve etc

Powered by blists - more mailing lists

Confused about mailing lists and their use? Read about mailing lists on Wikipedia and check out these guidelines on proper formatting of your messages.