<feed xmlns='http://www.w3.org/2005/Atom'>
<title>musl/src/regex, branch master</title>
<subtitle>musl - an implementation of the standard library for Linux-based systems</subtitle>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/'/>
<entry>
<title>fnmatch: do not match reversed ranges in bracket expressions</title>
<updated>2026-10-05T21:49:58+00:00</updated>
<author>
<name>Rich Felker</name>
<email>dalias@aerifal.cx</email>
</author>
<published>2026-10-05T21:49:58+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=b1efda5b91735e33635376ca11c9b497a1c66e39'/>
<id>b1efda5b91735e33635376ca11c9b497a1c66e39</id>
<content type='text'>
POSIX allows a reversed-order range to be treated as invalid (forcing
the bracket to be interpreted as a literal) or as valid but not
matching any characters. because we matched the starting character of
the range before identifying it as a range, however, [z-a] would match
'z'.

explicitly check for a subsequent non-terminal '-', and don't match if
one is present. the match, if any, will be caught by the range logic
in the next iteration.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
POSIX allows a reversed-order range to be treated as invalid (forcing
the bracket to be interpreted as a literal) or as valid but not
matching any characters. because we matched the starting character of
the range before identifying it as a range, however, [z-a] would match
'z'.

explicitly check for a subsequent non-terminal '-', and don't match if
one is present. the match, if any, will be caught by the range logic
in the next iteration.
</pre>
</div>
</content>
</entry>
<entry>
<title>fnmatch: fix processing bracket range with multibyte endpoint</title>
<updated>2026-10-05T21:44:14+00:00</updated>
<author>
<name>Rich Felker</name>
<email>dalias@aerifal.cx</email>
</author>
<published>2026-10-05T21:44:14+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=cb037c4630ddb87b48ad732bff04ebaa3d42148b'/>
<id>cb037c4630ddb87b48ad732bff04ebaa3d42148b</id>
<content type='text'>
the advancement logic was off-by-one, accounting for the p++ in the
for clause but not the '-' character. this resulted in attempting to
process the last byte of the range endpoint on the next iteration of
the loop.

for a single-byte endpoint this was harmless, but for a multi-byte
character, mbtowc would fail on the next iteration, preventing
matching of any subsequent part of the bracket.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
the advancement logic was off-by-one, accounting for the p++ in the
for clause but not the '-' character. this resulted in attempting to
process the last byte of the range endpoint on the next iteration of
the loop.

for a single-byte endpoint this was harmless, but for a multi-byte
character, mbtowc would fail on the next iteration, preventing
matching of any subsequent part of the bracket.
</pre>
</div>
</content>
</entry>
<entry>
<title>fnmatch: fix FNM_PERIOD failure of escaped '.' to match leading '.'</title>
<updated>2026-09-21T13:15:22+00:00</updated>
<author>
<name>Rich Felker</name>
<email>dalias@aerifal.cx</email>
</author>
<published>2026-09-21T13:15:22+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=c4e1bb3994c14ed5112c894d15a451bf00f0d501'/>
<id>c4e1bb3994c14ed5112c894d15a451bf00f0d501</id>
<content type='text'>
the new condition also allows forward progress into ordinary matching
if the pattern begins with a backslash. this is fine independent of
FNM_NOESCAPE and independent of what follows the backslash; if
FNM_NOESCAPE is active, the backslash is literal and will not match.
if FNM_NOESCAPE is not active, the pattern beginning with a backslash
ensures that the first character is literal.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
the new condition also allows forward progress into ordinary matching
if the pattern begins with a backslash. this is fine independent of
FNM_NOESCAPE and independent of what follows the backslash; if
FNM_NOESCAPE is active, the backslash is literal and will not match.
if FNM_NOESCAPE is not active, the pattern beginning with a backslash
ensures that the first character is literal.
</pre>
</div>
</content>
</entry>
<entry>
<title>fix integer overflow in gai_strerror, hstrerror and regerror</title>
<updated>2026-09-08T00:32:07+00:00</updated>
<author>
<name>Luca Kellermann</name>
<email>mailto.luca.kellermann@gmail.com</email>
</author>
<published>2026-05-10T04:34:20+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=c5e2f7e7ad36ea839fce45af8370cb3e9ddf0cda'/>
<id>c5e2f7e7ad36ea839fce45af8370cb3e9ddf0cda</id>
<content type='text'>
at least gai_strerror() and regerror() are specified to accept any int
value. if the value was close to INT_MAX (for gai_strerror()) or
INT_MIN (for hstrerror() and regerror()) a signed integer overflow
would occur.

fix this by converting the int argument to unsigned before doing
arithmetic.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
at least gai_strerror() and regerror() are specified to accept any int
value. if the value was close to INT_MAX (for gai_strerror()) or
INT_MIN (for hstrerror() and regerror()) a signed integer overflow
would occur.

fix this by converting the int argument to unsigned before doing
arithmetic.
</pre>
</div>
</content>
</entry>
<entry>
<title>regex: reject invalid \digit back reference in BRE</title>
<updated>2026-03-30T19:59:35+00:00</updated>
<author>
<name>Szabolcs Nagy</name>
<email>nsz@port70.net</email>
</author>
<published>2026-03-23T17:33:20+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=40acb04b2c1291f7d3091c61080109da11eea48b'/>
<id>40acb04b2c1291f7d3091c61080109da11eea48b</id>
<content type='text'>
in BRE \n matches the nth subexpression, but regcomp did not check if
the nth subexpression was complete or not, only that there were more
subexpressions overall than the largest backref.

fix regcomp to error if the referenced subexpression is incomplete.
the bug could cause an infinite loop in regexec:

 regcomp(&amp;re, "\\(^a*\\1\\)*", 0);
 regexec(&amp;re, "aa", 0, 0, 0);

since BRE has backreferences, any application accepting a BRE from
untrusted sources is already vulnerable to an attacker-controlled
near-infinite (exponential-time) loop, but this particular case where
the loop is actually infinite can and should be avoided.

ERE is not affected since the language an ERE describes is actually
regular.

Reported-by: Simon Resch &lt;simon.resch@code-intelligence.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
in BRE \n matches the nth subexpression, but regcomp did not check if
the nth subexpression was complete or not, only that there were more
subexpressions overall than the largest backref.

fix regcomp to error if the referenced subexpression is incomplete.
the bug could cause an infinite loop in regexec:

 regcomp(&amp;re, "\\(^a*\\1\\)*", 0);
 regexec(&amp;re, "aa", 0, 0, 0);

since BRE has backreferences, any application accepting a BRE from
untrusted sources is already vulnerable to an attacker-controlled
near-infinite (exponential-time) loop, but this particular case where
the loop is actually infinite can and should be avoided.

ERE is not affected since the language an ERE describes is actually
regular.

Reported-by: Simon Resch &lt;simon.resch@code-intelligence.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>glob: fix wrong return code when aborting before any matches</title>
<updated>2023-08-24T16:54:51+00:00</updated>
<author>
<name>Rich Felker</name>
<email>dalias@aerifal.cx</email>
</author>
<published>2023-08-24T16:54:51+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=79bdacff83a6bd5b70ff5ae5eb8b6de82c2f7c30'/>
<id>79bdacff83a6bd5b70ff5ae5eb8b6de82c2f7c30</id>
<content type='text'>
when the result count was zero, glob was ignoring a possible
GLOB_ABORTED error code and returning GLOB_NOMATCH. whether this
happened could be nondeterministic and dependent on the order of
dirent enumeration, in cases where multiple matches were present and
only some produced errors.

caught by Tor's test_util_glob.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
when the result count was zero, glob was ignoring a possible
GLOB_ABORTED error code and returning GLOB_NOMATCH. whether this
happened could be nondeterministic and dependent on the order of
dirent enumeration, in cases where multiple matches were present and
only some produced errors.

caught by Tor's test_util_glob.
</pre>
</div>
</content>
</entry>
<entry>
<title>remove LFS64 symbol aliases; replace with dynamic linker remapping</title>
<updated>2022-10-19T18:01:31+00:00</updated>
<author>
<name>Rich Felker</name>
<email>dalias@aerifal.cx</email>
</author>
<published>2022-09-26T21:14:18+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=246f1c811448f37a44b41cd8df8d0ef9736d95f4'/>
<id>246f1c811448f37a44b41cd8df8d0ef9736d95f4</id>
<content type='text'>
originally the namespace-infringing "large file support" interfaces
were included as part of glibc-ABI-compat, with the intent that they
not be used for linking, since our off_t is and always has been
unconditionally 64-bit and since we usually do not aim to support
nonstandard interfaces when there is an equivalent standard interface.

unfortunately, having the symbols present and available for linking
caused configure scripts to detect them and attempt to use them
without declarations, producing all the expected ill effects that
entails.

as a result, commit 2dd8d5e1b8ba1118ff1782e96545cb8a2318592c was made
to prevent this, using macros to redirect the LFS64 names to the
standard names, conditional on _GNU_SOURCE or _LARGEFILE64_SOURCE.
however, this has turned out to be a source of further problems,
especially since g++ defines _GNU_SOURCE by default. in particular,
the presence of these names as macros breaks a lot of valid code.

this commit removes all the LFS64 symbols and replaces them with a
mechanism in the dynamic linker symbol lookup failure path to retry
with the spurious "64" removed from the symbol name. in the future,
if/when the rest of glibc-ABI-compat is moved out of libc, this can be
removed.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
originally the namespace-infringing "large file support" interfaces
were included as part of glibc-ABI-compat, with the intent that they
not be used for linking, since our off_t is and always has been
unconditionally 64-bit and since we usually do not aim to support
nonstandard interfaces when there is an equivalent standard interface.

unfortunately, having the symbols present and available for linking
caused configure scripts to detect them and attempt to use them
without declarations, producing all the expected ill effects that
entails.

as a result, commit 2dd8d5e1b8ba1118ff1782e96545cb8a2318592c was made
to prevent this, using macros to redirect the LFS64 names to the
standard names, conditional on _GNU_SOURCE or _LARGEFILE64_SOURCE.
however, this has turned out to be a source of further problems,
especially since g++ defines _GNU_SOURCE by default. in particular,
the presence of these names as macros breaks a lot of valid code.

this commit removes all the LFS64 symbols and replaces them with a
mechanism in the dynamic linker symbol lookup failure path to retry
with the spurious "64" removed from the symbol name. in the future,
if/when the rest of glibc-ABI-compat is moved out of libc, this can be
removed.
</pre>
</div>
</content>
</entry>
<entry>
<title>fix failure of glob to match broken symlinks under some conditions</title>
<updated>2019-08-07T06:38:45+00:00</updated>
<author>
<name>Rich Felker</name>
<email>dalias@aerifal.cx</email>
</author>
<published>2019-07-11T20:55:17+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=e408deefeb1a60b6e9e1bb63393590926f96ee64'/>
<id>e408deefeb1a60b6e9e1bb63393590926f96ee64</id>
<content type='text'>
when the pattern ended with one or more literal path components, or
when the GLOB_MARK flag was passed to request that glob flag directory
results and the type obtained by readdir was unknown or inconclusive
(symlink), the stat function was called to evaluate existence and/or
determine type. however, stat fails with ENOENT for broken symlinks,
and this caused the match to be omitted from the results.

instead, use stat only for the unknown/inconclusive cases with
GLOB_MARK, and otherwise, or if stat fails, use lstat existence still
needs to be determined. this minimizes the number of costly syscalls,
performing both only in the case where GLOB_MARK is in use and there
is a final literal path component which is a broken symlink.

based on/simplified from patch by James Y Knight.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
when the pattern ended with one or more literal path components, or
when the GLOB_MARK flag was passed to request that glob flag directory
results and the type obtained by readdir was unknown or inconclusive
(symlink), the stat function was called to evaluate existence and/or
determine type. however, stat fails with ENOENT for broken symlinks,
and this caused the match to be omitted from the results.

instead, use stat only for the unknown/inconclusive cases with
GLOB_MARK, and otherwise, or if stat fails, use lstat existence still
needs to be determined. this minimizes the number of costly syscalls,
performing both only in the case where GLOB_MARK is in use and there
is a final literal path component which is a broken symlink.

based on/simplified from patch by James Y Knight.
</pre>
</div>
</content>
</entry>
<entry>
<title>glob: implement GLOB_TILDE and GLOB_TILDE_CHECK</title>
<updated>2019-08-06T18:03:31+00:00</updated>
<author>
<name>Ismael Luceno</name>
<email>ismael@iodev.co.uk</email>
</author>
<published>2019-07-25T21:50:48+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=49eacf29d29fa704b2feda879978ae64a825358c'/>
<id>49eacf29d29fa704b2feda879978ae64a825358c</id>
<content type='text'>
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
</pre>
</div>
</content>
</entry>
<entry>
<title>allow escaped path-separator slashes in glob</title>
<updated>2018-10-13T05:28:07+00:00</updated>
<author>
<name>Rich Felker</name>
<email>dalias@aerifal.cx</email>
</author>
<published>2018-10-13T04:55:48+00:00</published>
<link rel='alternate' type='text/html' href='http://git.musl-libc.org/cgit/musl/commit/?id=481006fd8887b80c4f794085179047e28102b01a'/>
<id>481006fd8887b80c4f794085179047e28102b01a</id>
<content type='text'>
previously (before and after rewrite), spurious escaping of path
separators as \/ was not treated the same as /, but rather got split
as an unpaired \ at the end of the fnmatch pattern and an unescaped /,
resulting in a mismatch/error.

for the case of \/ as part of the maximal literal prefix, remove the
explicit rejection of it and move the handling of / below escape
processing.

for the case of \/ after a proper glob pattern, it's hard to parse the
pattern, so don't. instead cheat and count repetitions of \ prior to
the already-found / character. if there are an odd number, the last is
escaping the /, so back up the split position by one. now the
char clobbered by null termination is variable, so save it and restore
as needed.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
previously (before and after rewrite), spurious escaping of path
separators as \/ was not treated the same as /, but rather got split
as an unpaired \ at the end of the fnmatch pattern and an unescaped /,
resulting in a mismatch/error.

for the case of \/ as part of the maximal literal prefix, remove the
explicit rejection of it and move the handling of / below escape
processing.

for the case of \/ after a proper glob pattern, it's hard to parse the
pattern, so don't. instead cheat and count repetitions of \ prior to
the already-found / character. if there are an odd number, the last is
escaping the /, so back up the split position by one. now the
char clobbered by null termination is variable, so save it and restore
as needed.
</pre>
</div>
</content>
</entry>
</feed>
