Commit graph

722 commits

Author SHA1 Message Date
Michael Giacomelli
9477e371a3 wma: fix overflow in the LSP exponent curve
Very low bitrate WMA files send their spectral envelope as LSP
coefficients. The curve computed from them squares two products
whose value can reach about a million, in 16.16 fixed point. Where
the envelope should be lowest the squares overflowed, and those
points came out near the curve's maximum instead: bands at the top
of the spectrum up to 17 dB too loud, and everything else a few dB
low, as levels are taken relative to the maximum.

Sum the squares in 64 bits, and let pow_m1_4() take a value above
the 16.16 range.

Checked with perfsim (Sansa e200v1 build). The curve is within
0.2 dB of ffmpeg's formula at all 57600 points of 80 blocks traced
from two files. Against ffmpeg's decode, as the mean difference in
band levels over the whole file, with the noise coding fix applied
before and after:

  StutteringFile (22 kHz mono, 20 kbps)    1.6 dB -> 0.3 dB
  shoujicomcn2117164 (same)                0.2 dB -> 0.1 dB
  terra8k (11 kHz mono, 8 kbps)            0.3 dB -> 0.0 dB

Only the five files that use LSP exponents change; the other 23
are byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 20:47:12 -04:00
Michael Giacomelli
11b12894ad wma: fix the level of noise coded high bands
Low bitrate WMA files code their upper bands as noise at a given
power. Those bands came out 12 to 25 dB too quiet, which dulled
the treble. Two things were wrong.

The exponent pointer was moved too far in blocks whose exponents
have another resolution. This is ffmpeg's r20756 (f78501b264,
"Fix apparent 10l typos introduced in r8627"); r8627 was merged
here in 2f1da8d24a but the fix for it never was.

The fixed point gain of a band lost nearly all its precision: it
was zero for three bands in four in the file traced, and the sum
for a band's power overflowed in a quarter of them. Compute the
gain in 64 bits, once a band. The factor for the noise below the
first coded coefficient (WMA v1 only) was 16 bits too large and is
corrected by the same reasoning; no file here exercises it.

Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode,
as the mean difference in band levels over the whole file:

  01 - Jane Austen (44.1 kHz, 32 kbps)   5.5 dB -> 0.0 dB
  07.Devil.In.My.Mind (same)             6.2 dB -> 0.0 dB
  beyonthepain907z (same)                6.8 dB -> 0.0 dB
  moshimoashitaga (same)                 2.8 dB -> 0.0 dB

Four more noise coded files, already within 0.3 dB, now match too.
The output of the 15 files without noise coding is byte-identical.
Files that use LSP exponents still differ in their top bands; that
is a separate fault in the LSP curve.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 20:47:12 -04:00
Michael Giacomelli
fb2ee6c08f wma: decode every codec packet in an ASF payload
A payload normally holds one codec packet of blockalign bytes, but
some files put several in each. Only the first was decoded, so such
a file played one packet in every 8 or 15, as a few seconds of
broken sound, and then ended.

Step through the payload in blockalign sized packets.

Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode:

  doors-test.wma (WMA v2, 11 kHz mono), 15 packets a payload:
    80896 of 1205760 samples before, all of them after, 86 dB SNR
  test.wma (WMA v1, 44.1 kHz stereo), 8 packets a payload:
    1294336 of 10346496 samples before, all after, 114 dB SNR

The output of 26 other WMA files is byte-identical before and after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 20:47:12 -04:00
Michael Giacomelli
6ffbee954b wmapro: fix the level and accuracy of 16 bit streams
A WMA Professional stream of 16 bits per sample played about 48 dB
too quiet and with a noise floor near -70 dBFS.

ffmpeg scales the transform's output by the stream's sample size.
That was dropped when the decoder was converted to fixed point
(d884af2b99, 16284ae8ae), and the output is passed to the DSP as if
every stream had 24 bits. A stream of fewer bits has a lower
quantization step to match, so it came out low by the difference,
and it used the bottom of the integer quantization table, where the
factors have only a few significant bits.

Decode a 16 or 20 bit stream at the level of a 24 bit stream: use
that stream's quantization step, and scale each band's factor by
the ratio that is left. 24 bit streams are not affected.

Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode
of a 16 bit, 192 kbps stereo file, the only such file to hand:

            level      SNR vs ffmpeg   noise
  before    -48 dB     37 dB           -70 dBFS
  after     correct    85 dB           -117 dBFS

A 24 bit file's output is byte-identical before and after. The
20 bit case follows the same rule but is untested.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 20:13:48 -04:00
Michael Giacomelli
4502948cd8 wmapro: add a two-channel path to the channel transform
inverse_channel_transform() ran its general N-channel matrix loop
for every sample of a stereo stream, where the matrix is exactly
+-1.0. That loop was a quarter to a third of the whole decode, and
GCC 9.5.0 compiles it worse than 4.9.4 did, which made the codec
6-9% slower on ARM7TDMI after the toolchain update.

Handle a group of two channels separately: add and subtract when
the matrix is +-1.0, and a plain four multiply loop otherwise. More
than two channels still use the general loop.

This reverses the regression from the GCC 9.5.0 update and goes
well past it. The loop the newer compiler handled badly is no longer
used for stereo, so the two compilers now give the same speed to
within 1%, about 25% faster than the codec was with GCC 4.9.4
(estimated with perfsim, e200v1, wmapro_141k: 25.81 MHz with 4.9.4
before this change, 19.5 MHz with either compiler after it).

Output is bit-identical: whole-file PCM hashes match before and
after for five stereo files at 55-271 kbps, built with GCC 9.5.0
and with 4.9.4, and also with the multiply path forced on.

Measured with test_codec, wmapro_141k.wma, MHz for real time:

  Sansa e200v1  27.99 -> 19.70
  Sansa Clip+   21.78 -> 15.80

Estimated with perfsim for the other files (e200v1 / Clip+):

  wmapro_55k   25.21 -> 17.06 / 20.02 -> 13.71
  wmapro_80k   26.17 -> 18.01 / 20.75 -> 14.44
  wmapro_173k  28.52 -> 20.29 / 22.55 -> 16.25
  wmapro_271k  30.85 -> 22.52 / 24.34 -> 18.04

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 19:29:19 -04:00
Michael Giacomelli
374f7c0334 flac: ARM assembly for the 64-bit LPC filter
Replaces flac_lpc_32_c, the wide predictor path every 24-bit stream
takes, and 16-bit streams encoded at high coefficient precision. The
coefficients are invariant for the whole call, but the C loop reloaded
all of them for every output sample and spilled its loop bound to the
stack on top of that. Orders 1-8 now keep every coefficient in a
register; orders 10 and 12, which is what -8 emits, get their own
unrolled loops instead of the generic chunked one. An ARMv5E kernel
using the packed 16-bit multiplies is included behind
FLAC_LPC32_NARROW_ASM, off by default: it needs bps <= 16, and only
about half the subframes of such a stream can use it.

Measured with test_codec on 24-bit/96kHz streams. At predictor order 12:
36.24 MHz on e200 (ARMv4) against 62.29 MHz before. The same file
improves from 37.65 MHz to 27.3 MHz on Clip+ (ARMv5).

Bit-exact against flac -d over four complete streams on both targets,
with and without the assembly, and the kernels are checked against an
int64_t reference across every order 1-32, qlevel 0-15 and coefficient
precision 1-15.

The rarely-executed orders stay in DRAM: the codec's IRAM window on
PP502x is nearly full and demoting them measured no cost, leaving 96
bytes free. FLAC_LPC32_NO_IRAM demotes the whole filter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Icd8bd298973b9f5c26c103af8adf47b03489d69f
2026-10-05 12:19:09 -04:00
Michael Giacomelli
a6b7321d9f libdemac: reciprocal table for the pivot division on ARMv5+
Of the two divisions a sample left in the entropy decoder, the first
divides the range by the Rice pivot, which is small: below 1024 for
99.9% of the samples of a 16-bit test file and never as much as 4096.
Keep a table of reciprocals, filled in as divisors turn up, and
divide with a 32x32->64 multiply and one correction.  Larger divisors
go to the division routine as before.  The other division, by help,
has no such pattern.

This is for ARMv5 and later without a hardware divide.  ARMv4 has
its own divider with a reciprocal table already, and a long multiply
is slow there; its codec is unchanged.

The table is 16 KB of bss.  With r = (2^32 - 1) / n the estimate is
the quotient or one less for any 32-bit numerator, checked against
true division for every n below 4096.

MHz for real time in perfsim's model of the Clip+, -c1000, -c2000
and -c3000: 32.5, 48.4, 76.4 to 30.1, 46.0, 74.0.  A Clip+ measures
30.63 at -c1000, from 32.85, with the same checksum, 1008ffab.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I555ff83d7e19919848c8fb7b8d12ecfa671a976b
2026-10-05 10:19:38 -04:00
Michael Giacomelli
aa7ebbe057 libdemac: entropy state in registers, one less divide
After the division routines the entropy decoder was the largest part
of Monkey's Audio at the fast levels, with two things to gain.

All of its state, the range coder's and the Rice parameters, was in
statics, so each use was a load and each update a store: about half
of the function's time on ARM7TDMI, where a load is 3 cycles.  Copy
it to locals for the length of a block and write it back after.  The
functions that take a pointer to it have to be inlined for that to
work, and on ARMv4 gcc left the per-sample one out of line, so force
them.

The symbol was found by dividing low by help and searching the count
table for the quotient.  counts[n] <= low / help is the same as
counts[n] * help <= low, so search with the multiply instead: the
first symbols are by far the likeliest.  That leaves two divisions a
sample from three.

MHz for real time in perfsim's models, -c1000, -c2000 and -c3000:
  Clip+ (ARMv5)  37.9, 53.8, 81.8  to  32.5, 48.4, 76.4
  e200 (ARMv4)   45.3, 68.6, 111.4 to  35.2, 58.5, 101.3
A Clip+ measures 32.85 at -c1000 (61.4 before this and the division
change), with test_codec's checksum, 1008ffab, the same as the old
code's.  The standalone decoder gives identical output for all four
levels tested.  The path for files older than 3.98 has the same
change and was not tested, for want of a file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Idf9002a5b5f3a423f44f70443936ea5ac427cd12
2026-10-05 10:19:38 -04:00
Michael Giacomelli
5d8ce62b45 codecs: link libarm_support for its division routines
lib/arm_support/support-arm.S was written to replace libgcc's
division for ARM, and bff5a35c3c (FS#10943, 2010) added it to the
core, the plugin library and the codec library alike.  When
1501df045f (2013) replaced EXTRA_LIBS with explicit lists, plugins
kept it and codecs did not, and they have taken their division from
libgcc since.

With the gcc 4.4 toolchain that cost little: its libgcc had a
routine that used clz.  With gcc 9.5 libgcc has no soft-float ARMv5
variant, so ARMv5 targets get the ARMv4 routine, a shift and
subtract loop of about 130 cycles a division.

Monkey's Audio divides two or three times a sample in its range
decoder and, without codec IRAM, does it in C.  On a Clip+ it is
11% to 21% slower than 3.14 was, with over half of -c1000's decode
in __udivsi3.  MP2 is 4% to 7% slower.

Put libarm_support back, ahead of libgcc.  In perfsim's model of
the Clip+ a division falls to 44 cycles, Monkey's Audio by 38%, 30%
and 22% at -c1000, -c2000 and -c3000 (60.9 to 37.9 MHz at -c1000),
and MP2 by 3% to 7%; nothing else moves by more than 0.6%.  A Clip+
decodes -c1000 with the same checksum as before.

A division by zero in a codec goes to __div0 again, and so to the
firmware's handler, as it did with support-arm.S and with the old
libgcc (3.14's ape.codec calls it).  The gcc 9.5 libgcc returns
from its own stub instead.

Change-Id: I6a9a79ce8e85bca69870349a5c0823f392a578b6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 07:46:37 -04:00
Michael Giacomelli
0ed3734e5e flac: cope with frames larger than one buffer request
The frame decoder reads from one flat buffer and cannot refill it, but
request_buffer() only guarantees 32KiB of contiguous data (the buffering
guard area). Frames that can be larger than that, such as high
resolution or poorly compressible streams (FLAC decoder testbench file
31), could be handed to the decoder truncated, which read past the end
of the data and lost sync.

When a request returns less than the largest frame the stream can
contain (STREAMINFO max framesize, or a bound from block size, channels
and bit depth) and it is not the end of the file, copy the frame into a
private static buffer and decode from that. Streams whose frames always
fit never touch the buffer and pay one comparison per frame. The buffer
is 64KiB, or sized for the 4608 sample blocks of memory limited targets,
and is left out entirely when MEMORYSIZE is 2MB or less.

Also fail with a codec error, instead of advancing past the data, if a
decoded frame consumed more bytes than were provided.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: Ib4cf85ad513f8e096f73d13b4374d072e1745cb2
2026-10-03 18:17:56 -04:00
Michael Giacomelli
e45936397e opus: filter the PLC excitation in place
At a CELT to hybrid mode switch, opus_decode_frame() runs CELT's pitch
PLC nested in the new frame, and celt_decode_lost() held a copy of up to
2 KB of excitation for celt_fir().  On stackOverflow.opus that overran
the 9 KB codec stack on native targets such as the e200v2.  Filtering in
place from the last sample down needs no copy; output is bit-identical.

Worst case over all 242 mode switches of stackOverflow.opus, measured
under qemu: 9140 -> 7676 bytes below opus_decode(), against 8912 left
for it on the e200v2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Iaa9f57182ad86ab76c81c55bea72aa517cc18230
2026-09-30 11:25:30 -04:00
Michael Giacomelli
330236911b flac: decode residuals that use all 32 bits
The folded Rice value was unfolded with a signed shift, which is wrong
once it reaches 2^31, and the unary length limit was (INT_MAX >> k) + 2,
about half of what a 32-bit residual can need. Streams with very large
residuals (FLAC decoder testbench file 63) were misparsed, overran the
frame and lost sync at the next one. Unfold as unsigned and derive the
limit from UINT_MAX, clamped to INT_MAX. The fast path is unchanged.

The existing 0x80000000 error check now also works as intended, since
the Golomb reader's error value maps to it. The FLAC spec forbids a
residual of -2^31.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: I0bdce9db69e1b30565ecc23126923ba401b7deca
2026-09-30 08:24:15 -04:00
Michael Giacomelli
deaef503bb opus: ARMv5E SILK resampler interpolation
On ARMv5E and later the fixed-phase FIR cycles at 8, 12 and 16 kHz load
two samples a word and take two coefficients a word from a literal
pool; smla<x><y> picks the halves, so an odd-aligned window costs
nothing.  Outputs are paired as two interleaved accumulator chains.  The
FIR buffer is now word aligned.  Generated by
silk/arm/gen_resampler_armv5e.py.  Bit-exact; OPUS_ARM_NO_SILK_ASM
disables it, and config.h sets that on M-profile cores.

Measured on the Clip+, silk_5.opus (WB SILK):
20.93 -> 15.84 MHz, -24.3%.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Ibec71fd52542f768f22e2970c9f8c45118c708b9
2026-09-30 08:21:46 -04:00
Michael Giacomelli
15eec82d8f opus: ARMv4 SILK resampler interpolation
On ARMv4 the fixed-phase FIR cycles at 8, 12 and 16 kHz run as adds of
shifted samples rather than multiplies: every coefficient is a constant,
and a shifted add is one cycle where mla plus loading the coefficient is
five or six.  Each sample is loaded once per cycle and added into the
two or three outputs it feeds, sharing partial products such as 31x
between them, about 27 adds per output.  The kernels are generated by
silk/arm/gen_resampler_armv4.py.  Bit-exact; OPUS_ARM_NO_SILK_ASM
disables them.

Measured on the e200v1, silk_5.opus (WB SILK), with the previous commit:
36.31 -> 28.85 MHz, -20.5%.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Ic1fd5777e91d182248f43954e4b6ae977309dc5d
2026-09-30 08:21:46 -04:00
Michael Giacomelli
b6b2d95309 opus: fixed-phase SILK resampler interpolation
The 8, 12 and 16 kHz to 48 kHz steps visit only two or three FIR phases
in a fixed cycle, so each set of input samples is read once and reused
across outputs.  Bit-exact; OPUS_NO_SILK_FIXED_PHASE disables it.

Measured with silk_5.opus (WB SILK):
e200v1: 36.31 -> 33.38 MHz, -8.1%
Clip+:  21.66 -> 20.93 MHz, -3.4%

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Icb0f8ded62739dfee1f574e4888a54c778aa3d53
2026-09-30 07:54:43 -04:00
Michael Giacomelli
04ecb3c645 flac: don't divide by zero when STREAMINFO has no sample count
A total sample count of 0 means unknown. It made the track length 0 and the
bitrate estimate in flac_init() divided by it, crashing the codec (FLAC
decoder testbench file 45). Report a bitrate of 0 in that case.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-28 16:35:14 -04:00
Michael Giacomelli
a81fd7b77f flac: handle zero-width escaped rice partitions
An escape code with a raw bit width of 0 means every residual in the
partition is zero. get_sbits(&gb, 0) shifts by 32, which is undefined and
returned stale cache bits instead of 0, so such streams decoded to garbage
(FLAC decoder testbench file 64).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-28 16:35:14 -04:00
Michael Giacomelli
8f38274e90 warble: zero the mp3entry before reading metadata
print_mp3entry() dereferences mb_track_id, but get_metadata() does not set
every field of the uninitialized stack struct. For FLAC files this left
garbage in the pointer and warble segfaulted in strlen about a third of the
time, before decoding started. Clear the struct first.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-28 16:35:14 -04:00
Michael Giacomelli
c47b2e1be2 opus: name the ARM inline-asm gates after the cores they cover
Upstream's OPUS_ARM_INLINE_ASM means any ARM with inline assembly, with
OPUS_ARM_INLINE_EDSP layered on top for ARMv5E.  config.h instead defines
exactly one of them per core, so the names read as broader than they are,
and code added ARM_ARCH tests beside them to pin the scope down.  Renamed
to OPUS_ARM_ASM_ARMV4_ONLY and OPUS_ARM_ASM_ARMV5E_AND_LATER throughout
celt and silk, upstream files included; README.rockbox records it for the
next libopus sync.

No code change: opus.elf disassembly and section sizes are identical
before and after on ARMv4 (e200v1), ARMv5E (Clip+) and ARMv6 (iPod Nano
4G).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I89932d348dedd76748a1bc7b3f2e7de1c5be49c8
2026-09-28 12:10:57 -04:00
Michael Giacomelli
b12ef5e6f8 opus: gate the ARMv5E kernels for ARMv6 too
Four `ARM_ARCH == 5` checks -- in SOURCES and celt/arm/{bands_arm,
comb_filter_arm,vq_arm}.h -- excluded ARMv6 from every ARMv5E kernel:
denorm_band, haar1, comb_filter_const, celt_sat, deemphasis_stereo_simple,
exp_rotation1 and normres_scale all silently fell back to plain C on
ARM1136/ARM1176, since ARMv6 is a strict superset of the EDSP instructions
those kernels use. Widened to ARM_ARCH >= 5, matching config.h's own
OPUS_ARM_INLINE_EDSP gate, which was already ARM_ARCH > 4.

These four can't be fixed at the commits that introduced them: those
commits are already merged into master under different SHAs. A fifth
instance of the same bug, in celt/arm/mdct_armv5e.h, was fixed at its
origin commit instead, since that one is still open for review.

Verified on both native ARMv6 targets (iPod Nano 4G, ARM1176JZ-S; Gigabeat
S, ARM1136JF-S) and the hosted Samsung YP-R0 (ARM1176JZ-S, cross toolchain
built for the occasion): all three now link and call all 16 ARMv5E
kernels, where they linked and called zero before this fix. Decoded PCM
bit-identical to the ARMv5E build under qemu.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: I7e6f6b421b9c0203812fb45d2258c7eb81738d98
2026-09-27 05:52:23 -04:00
Michael Giacomelli
5a6935047e opus: hold cwrsi's row pair, and test it with one comparison
53% of the dimensions a decode walks hold no pulses, and that arm leaves
_k alone, so the two CELT_PVQ_U_ROW pointers stay valid.  U is
non-decreasing in _k, so p <= _i < q is the single unsigned test
(_i-p) < (q-p).  Bit-exact over 8.2M samples.

Modelled: -0.40% ARMv4, -0.78% ARMv5E; the function -4.8% and -7.1%.
Measured: e200v1 39.19 -> 38.96 MHz, Clip+ 28.13 -> 28.08 MHz.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id50be222e101aeabd9a1f4cf5380d3244da610ff
2026-09-27 05:37:13 -04:00
Michael Giacomelli
93d564426b opus: ARMv4 exp_rotation1
At stride 1 the rotation chain writes X[i+stride] and reads it straight
back, so one load an iteration is redundant and one store is dead.  The
kernel carries that value, narrowed, since mul reads all 32 bits where
smulbb does not.  Bit-exact over 8.2M samples.

Modelled: -0.87% ARMv4, the function -10.0%.
Measured: e200v1 39.59 -> 39.19 MHz.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Idbd1a088bc805fecfb9ee54fc372ad16a7d91606
2026-09-27 05:35:00 -04:00
Michael Giacomelli
a51adaac4f opus: place the hot decode path in IRAM on PP5022/5024
34 functions and the ARMv4 kernels, chosen by a greedy fill of the free
codec IRAM window ranked by cycles per byte.  The mixed-radix
butterflies give their IRAM back, being unreachable under the prime
factor transform.

The cycle model does not see this at all: it models no instruction cache.
Measured: e200v1 48.96 -> 42.60 MHz; reclaiming the mixed-radix IRAM was
a further -0.52%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I9e22025cba99521750a3b664cd0d59357fcd0810
2026-09-26 17:48:52 -04:00
Solomon Peachy
40fdf6ac11 libopus: Fix yellow and red in da9df96c30
* simulator warnings
 * non-arm warnings
 * armv7-m errors
 * armv7-a errors

Change-Id: I285ab466c07170e632f1bec488eb471bcbd4a566
2026-09-26 14:49:52 -04:00
Michael Giacomelli
da9df96c30 opus: Good-Thomas FFT for the backward MDCT
Every 48 kHz CELT length is 15 times a power of two and the factors are
co-prime, so the inter-stage twiddles -- 73% of the FFT multiplies at
N=480 -- vanish.  New pfa_fft15 and mdct_postrot_pfa kernels on both
cores; accuracy also improves 0.3 to 0.4 dB against opusdec.

Modelled: -6.17% ARMv4, -1.82% ARMv5E.
Measured: e200v1 42.33 -> 40.10 MHz, Clip+ 29.30 -> 28.24 MHz.
Rescheduling these kernels, folded in here, measured a further -0.75% on
e200v1 and -0.39% on Clip+.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ifa4a30d1045e905522828958949954615faa440d
2026-09-26 13:57:35 -04:00
Michael Giacomelli
7b4d1a7f75 opus: dual_inner_prod on ARMv5E
One ldr fetches two celt_norm coefficients and smlabb/smlatt take the
halves apart; scalar fallback when the three pointers disagree on
alignment.

Modelled: -0.88% ARMv5E.
Measured: Clip+ 28.13 -> 28.08 MHz, the same build with and without it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I79719132ccecfb7664616bca019e3856ee3479da
2026-09-26 13:26:45 -04:00
Michael Giacomelli
ff762858c7 opus: build the decoder without the encoder halves
Rockbox lists no encoder, so CELT_DECODE_ONLY folds away the thirteen
encoder branches in celt/bands.c and stops gcc keeping their values live
across quant_partition's recursive calls.  2,976 bytes smaller.

Modelled: -0.47% ARMv4, -0.69% ARMv5E.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ie96a0a63e035ab0cdcdbe9eeeb6691f24fecfa41
2026-09-24 17:47:58 -04:00
Michael Giacomelli
07557038b4 opus: tabulate bitexact_log2tan on PP5022/5024
Its reachable input set is 2,985 (qn,i) pairs for any stream ever, so it
tabulates exactly in 6,484 bytes.  Verified exhaustively against the
compiled functions.  Only built for PP5022/PP5024, where the tables fit
the 80 KB IRAM window; elsewhere it measured no gain.

Measured, e200v1: 39.18 -> 38.96 MHz with the tables in IRAM, 39.07 with
them in DRAM.  Clip+: 28.05 -> 28.08 MHz, so not built there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I38417793dc337ec6cb61139cc919d23ac09b04dd
2026-09-24 17:47:33 -04:00
Solomon Peachy
0918a068eb opus: More surgical update to 9a972e7f51
This re-enables upstream ASM optimizations for Cortex-M.  Only the
our downstream improvements are disabled, as they do not assemble in
thumb2 mode.

Change-Id: Icca1b3dbf04786c7714fb4eef92ad66aa55132f3
2026-09-24 17:24:24 -04:00
Solomon Peachy
9a972e7f51 opus: Correct enablement of new ARMv5e optimizations
* Only use new ARMv5e optimizations on classic (non-M) profile
 * fix inconsistent ARM_ARCH >= 5 vs == 5

Fixes red in 0c4345475a and ae223933bf

Change-Id: Ieb688679d2d698a19870b1a63b4151a8593e04e0
2026-09-24 16:59:07 -04:00
Michael Giacomelli
ae223933bf opus: ARMv5E FFT butterflies
radix-3, 4 and 5.  The gain is bookkeeping: twiddles addressed by
displacement from one base register, post-indexed stores, and C_MUL's
Q15 doubling folded into the add that consumes it.

Modelled: -4.72% ARMv5E, opus_fft_impl -25.3%.
Measured: Clip+ 30.89 -> 29.33 MHz.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I173667ebb13308f83bd6fcb4e876f5a56c698cd3
2026-09-24 16:33:04 -04:00
Michael Giacomelli
0c4345475a opus: ARMv5E assembly for the backward MDCT inner loops
Pre-rotation, post-rotation and mirror, with a packed-twiddle complex
multiply throughout.

Modelled: -8.0% ARMv5E against the C loops.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I3fcaf2343e04e3f3f4462fc442f8fac623729631
2026-09-24 16:32:52 -04:00
Michael Giacomelli
08d3332edf opus: ARM stereo de-emphasis kernel
deemphasis_stereo_simple on both cores, with the filter state kept
unshifted and shifted inside the add that consumes it.

Modelled: -0.53% ARMv4, -0.96% ARMv5E.
Measured with the three preceding commits: e200v1 49.42 -> 48.96 MHz,
Clip+ 32.83 -> 30.89 MHz.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id09297974f3b66a5a193b524d68fd8844f71cc02
2026-09-24 15:50:19 -04:00
Michael Giacomelli
31f2a27f1e opus: inline EC_ILOG on ARMv4
ARMv4 has no CLZ, so all 8,680 ilog2 calls went through libgcc's
__clzsi2.  Fifteen branchless instructions replace it.

Modelled: -0.91% ARMv4; ARMv5E unaffected, it already emits CLZ.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ib84b97acb687da4101e02bad1b0e8bbd256d32df
2026-09-24 10:39:52 -04:00
Michael Giacomelli
8eb6b05945 opus: ARM comb filter and saturation kernels
comb_filter_const on both cores, and celt_synthesis's SIG_SAT clamp
four samples at a time through one ldm and one stm.

Modelled for the saturation loop: -1.05% ARMv4, -1.42% ARMv5E.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I381afab451eb29a91513ce6ab29c1b2d565b7ee9
2026-09-22 17:26:11 -04:00
Michael Giacomelli
f26c9557c2 opus: ARM kernels for the 16-bit band loops
denormalise_bands on both cores; exp_rotation1, haar1 and the
normalise_residual scaling loop on ARMv5E.  Bit-exact.

Modelled: -0.50% ARMv4, -2.91% ARMv5E.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: If40f5dbe7cb9e7c12aa0a5b4ac9e73433a850b6e
2026-09-22 17:20:32 -04:00
Michael Giacomelli
9e7b81f269 opus: exact bounded divide for the range decoder
98.6% of __udivsi3 calls are ec_decode and ec_decode_bin computing
val/ext, where the quotient only matters below 2^16.  Exact over 80
million cases including corrupt-stream values.

Modelled: -1.34% ARMv4, -2.37% ARMv5E.
Measured with the previous commit: e200v1 50.80 -> 49.42 MHz,
Clip+ 33.40 -> 32.83 MHz.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ia09cb6cd8e42d4885e4d5be0f84862b9b951136c
2026-09-22 16:34:14 -04:00
Michael Giacomelli
d514ee9282 opus: ARMv4 assembly for the backward MDCT inner loops
Cuts realtime decode on the Sansa e200v1 from 52.1 MHz to 50.8 MHz.
clt_mdct_backward was the largest remaining item at 13.5% of decode.

Only the three inner loops move to assembly.  The setup stays in C, so
mdct.c remains readable and the assembly needs no knowledge of
mdct_lookup.

What the compiled loops lose is registers.  Each needs more live values
than gcc can hold, so it spills the loop-invariant pointers, strides and
limits and reloads them every pass: five stack accesses per iteration in
the post-rotation alone.  Holding the twiddle as a 16-bit value and
accumulating the product pair with smull/smlal is what makes the
bookkeeping fit, needing seven live registers where the shifted
MULT16_32_Q15 form needs nine.

ldm/stm helps only where the addressing allows.  The post-rotation walks
the buffer from both ends and so reads and writes contiguous pairs.  The
pre-rotation reads the spectrum through a runtime stride and writes
through the bitrev table, so only its 8-byte output pair merges, and the
TDAC mirror merges nothing.

Over 160 ms of stereo music, traced under qemu:

  clt_mdct_backward  1,037,962 ->   900,982   -13.2%
  whole decode       7,695,876 -> 7,558,896    -1.8%
  loads                650,157 ->   611,667    -5.9%
  stores               350,605 ->   323,605    -7.7%
  multiplies           337,493 ->   337,493   unchanged

Accuracy improves substantially, because all three loops keep 32 bits of
each Q15 product where MULT16_32_Q15_armv4 drops the low bit, and the
backward MDCT applies three such rounds per sample.  The rounding SNR of
the backward transform rises about 9.5 dB, and its worst case error falls
from 708 to 186.  Decoded output differs from the previous build in 90 of
15,360 samples, each by one LSB.

Build with OPUS_ARM_NO_MDCT_ASM to select the C loops instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I3c4404b4dbe581d8bcf1f357266a658a068fdb50
2026-09-22 07:26:56 -04:00
Michael Giacomelli
b18f5d6d65 opus: ARMv4 assembly FFT butterflies, and an exact radix-5 rewrite
Cuts realtime Opus decode on the Sansa e200v1 (PP502x, ARM7TDMI) from
55.55 MHz to 52.1 MHz.

Two exact changes to celt/kiss_fft.c first:

 - kf_bfly5 folds the four cosine products into a shift and a single
   multiply.  cos(2*pi/5) + cos(4*pi/5) is exactly -1/2, and the Q15
   constants satisfy that identity exactly (10126 - 26510 == -16384), so
   the substitution gives up no accuracy.
 - kf_bfly3, kf_bfly4 and kf_bfly5 peel the pass that twiddles by
   twiddles[0], which is 1, by rotating the loop rather than duplicating
   the body.  In fixed point twiddles[0] is 32767 rather than 32768, so
   skipping it also drops a small systematic gain error.

Then celt/arm/kiss_fft_armv4_asm.S replaces the radix-3, radix-4 and
radix-5 bodies, reached through OVERRIDE_kf_bfly3/4/5.  The compiled
kernels spill their loop-invariant pointers and reload them every pass,
and gcc will not form ldm/stm from contiguous C accesses: it reorders the
loads while scheduling and does not hand out ascending register pairs.
The assembly keeps the bookkeeping resident and sends the transient
butterfly values to the frame instead, in bursts.

Over 160 ms of stereo music, traced under qemu and costed with an
ARM7TDMI model:

  FFT cycles    1,876,926 -> 1,430,748   -23.8%
  whole decode  8,142,054 -> 7,695,876    -5.5%
  loads           769,935 ->   650,157   -15.6%
  stores          409,999 ->   350,605   -14.5%
  multiplies      350,909 ->   337,493    -3.8%
  text              3,368 ->     2,656 bytes

The radix-5 fold accounts for the whole multiply reduction.  The assembly
leaves the count untouched and wins purely on memory traffic.

Accuracy improves by up to 2.9 dB rather than degrading, because the
assembly keeps all 32 bits of each Q15 product where MULT16_32_Q15_armv4
drops the low bit.  Decoded output differs from the C build in 27 of
15,360 samples, each by one LSB.

Build with OPUS_ARM_NO_FFT_ASM to select the C butterflies instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I2e8a0904da8a85b654ff648a01234348977a0edc
2026-09-22 07:26:29 -04:00
Solomon Peachy
305acca1f0 dsp: Add missing .type <symbol>, %function to arm asm
Several symbols were missing these annotations so the compiler wasn't
handling thumb interworking correctly.

Should resolve tone control crashes seen on clipv1 and our handful of
other 2MB armv5 targets that we build (mostly) in thumb mode.

Change-Id: If8c533a5f10b6592a2d2bec519f2c4d6539d7a66
2026-09-19 21:18:30 -04:00
Aidan MacDonald
3a57f2f721 Add "rbfs" prefix to native filesystem functions
Most of the churn here occurs because 'filesize' is one of
the redefined filesystem functions, but some of the structs
used by the native FS code also include a 'filesize' member
variable which gets renamed by the macro in some but not all
source files.

It's easier to rename 'filesize()' to 'ffilesize()' rather
than try to clean up the macro mess or renaming the struct
members.

There is weirdness with root_realpath() which now breaks on
native builds because it was assumed to be unprefixed there.
dir_get_info() was unprefixed everywhere but this just seems
inconsistent; make it follow the FS_PREFIX() convention too.

Change-Id: Ic3700c6234ea45f32679c1a8429d70fdb8f4088a
2026-09-17 16:22:31 -04:00
David Cormier
064c165367 libm4a: handle sparse chunk maps and video-first MP4
Size the reduced stco lookup with ceiling division. The table stores
chunk entries 0, divider, 2 * divider, and so on, so floor division
allocated one entry too few whenever the original count had a
remainder.

Treat a valid media-information atom whose first child is not smhd as
a non-audio track and skip the remainder. Continue scanning subsequent
tracks so AAC decoding works when an MP4 places its video track before
the audio track, while still rejecting malformed atom sizes and
malformed sound headers.

Keep these container fixes independent of the H.264 player and target
driver so they can be reviewed and applied to libm4a on their own.

Build-tested as part of the normal and isolated-runtime iPod 6G
configurations and hardware-tested with AAC audio in M4V playback.

Change-Id: I8c711525932f54ecbdad982c7f7ddc9490c5668d
2026-09-07 09:17:27 -04:00
Mauricio Garrido
e498c0171a 3ds: Port refactor.
This commit does the following changes to the 3ds port:

- Rename the target from ctru to 3ds.
- Rename all files and functions with the ctru naming convention to 3ds.
- Created a new file and folder structure that will better integrate future console ports that share the same codebase.
- Fixed a buffer overflow bug in pcm code.

Change-Id: I17c6f86df64eb99dd2b653485d70832ff46b2ba8
2026-09-01 08:39:07 -04:00
Michael Giacomelli
7fef95dc08 warble: fix build on targets with HAVE_RECORDING
Change-Id: Ic04392da41177f51dd5f392b288e078a76298839
2026-08-25 08:01:09 -04:00
Skye
e764656ab7 Allow customizing EQ filter types
Allows any EQ band to be set to any of Low Shelf, Peak, or High Shelf, instead of hardcoding the types per band.

Change-Id: I470ab916359092ba465e7b6331baed3bf11b2fc9
2026-07-30 07:48:45 -04:00
Roman Artiukhin
958e84c042 codec: flac: fix playback of certain files
Use flac_seek by time even when elapsedtime is 0, and apply it as a fallback for failed offset seeks since it provides more robust error recovery.

Change-Id: I438888fba02bda38137f3f1347bb1f657ef9c166
2026-07-23 20:27:57 +03:00
Solomon Peachy
6c4107d1ad libspeex: Silence a spurious warning under Clang
Change-Id: I9790e92dd9120afff83114925b8e6eb482d6e7df
2026-07-15 16:22:18 -04:00
Solomon Peachy
b9c7b0e910 fix: new yellow
for #pragma GCC diagnostic, GCC must be capitalized.

Change-Id: I1d760dd83b4dc29590454f4f4e09c5dface2c48f
2026-07-11 08:52:12 -04:00
Solomon Peachy
9ebe34d570 speex: Silence spurious build warning when building under rbutil
Change-Id: Ida23a960c54de1a46d0787795efb7b28b6427939
2026-07-11 08:04:10 -04:00
Solomon Peachy
ebf42dae68 skin_parser: Fix regression in 0d5afa6d
Typo in the #ifdef, accidently didn't get committed.

Change-Id: If74de478ed8f236b78e16cd3cb3957ecde3339b3
2026-07-05 19:12:25 -04:00