RFC 7845 specifies R128_TRACK_GAIN and R128_ALBUM_GAIN for loudness
normalization in Opus files, but Rockbox only understood the
REPLAYGAIN_* tags. Parse the R128 tags as Q7.8 dB values relative to
-23 LUFS and add 5 dB to match the ReplayGain reference level. If a
file has both kinds of tags, the R128 tags take priority.
Fixes FS#13624
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I0ef08a7897ec5ee8d1e5d99e7d9ea3901847fa52
afr_configure() returned 1 for every message. For
DSP_PROC_NEW_FORMAT that is PROC_NEW_FORMAT_DEACTIVATED, so the
DSP did not call the stage for the buffer that brought a new
format: the first buffer of each track went out unfiltered, and
the filters started on the second. That is 128 samples of a WAV
file, and the whole first block of a codec with 32-bit
non-interleaved output, 4608 samples of a FLAC file.
A codec sets the format again at the start of every track, so
between gapless tracks this put a short stretch of unfiltered
audio into filtered audio, heard as a click at the join.
Return 0, as the other stages do.
Found with test_codec's DSP checksums on a Sansa Clip+: a model
of the DSP with fatigue reduction left out of each file's first
buffer reproduces the sixteen checksums of two runs, and none of
them with it left in. Listening on the same player, the click
between gapless tracks is gone with this change.
Change-Id: I7a1193fa0a4ee0c983cbcc5c1af201e1678579d0
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What the built-in recording screen shows and the skin tags did not:
- %RS size of the current recording, as "1.5MB"
- %RP seconds in the pre-record buffer, empty unless pre-recording
- %Rc clips counted by the peak meter
- %Rt trigger state by name; as a conditional off, ready, steady, go,
postrec, retrig, continue
- %Rw recording warnings in hex, empty while there are none
- %Ri recording source by name; as a conditional the same on every
target: mic, line in, digital, FM radio
- %Rg gain of the source in dB, the left channel's for line in and FM
The manual gets a Recording section covering these and the recording
tags that were there, none of which it described.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I889f95fd472c472c2589393a6c77eff4a4fd5c8f
b7170e03c4 added window_lookup_mid, 512 bytes, to IRAM on every
target. On the targets with 48KB of IRAM for a codec that left
the ATRAC3 codecs 48 bytes too large to link (iPod Color and the
other PP502x players with 96KB of IRAM).
Put the table in IRAM only on the targets with more, like the
decoder's other large tables. Nothing changes there.
Built for iPod Color, where the two codecs now have 476 bytes of
IRAM to spare, and for the iriver H120.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The decoder gave the DSP samples with two fractional bits (sample
depth 17), far fewer than any other codec. The DSP's filters are
only as precise as the samples they are given: with an equalizer
band at a low frequency the output was noise, or held at full
scale.
Scale the last stage of the synthesis filter up by 11 bits and
set the sample depth to 28. On ARMv4 the bits are taken from the
64-bit sums of the dewindowing, which had them. Elsewhere the
input of that stage is scaled, in its matrixing step; that is not
done on ARMv4 because its multiplier takes longer for the larger
operands (51.97 against 48.85 MHz below). The Coldfire path uses
the C matrixing and has not been run.
Seven files decoded under perfsim for the Sansa e200v1 and Clip+:
shifted down again, the e200v1 output is within one step of the
old output, and the Clip+ output no further from the e200v1's
than before. Through a ten band equalizer with bands at 32 and
64 Hz (the filter fix of the DSP included), the noise of
atrac3_lp2_132.oma falls from -64 dB to -101 dB relative to the
signal.
Estimated with perfsim (not measured on a device),
atrac3_lp2_132.oma: e200v1 48.81 -> 48.85 MHz, Clip+ 26.91 ->
26.96 MHz.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The dewindowing loop is 72% of ATRAC3 decoding on ARM7TDMI. The
window is symmetric, so read only its first half: each pair of
coefficients serves two inputs from the front of the 48 and, with
the two swapped, two from the back. That halves the coefficient
loads and lets them be done four at a time.
The sums are kept in 64 bits, so the new order of the additions
does not change the result: the PCM output of seven files is
byte-identical.
Estimated with perfsim for the Sansa e200v1 (not measured on the
device), atrac3_lp2_132.oma: 53.74 -> 48.81 MHz.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ARMv5E dewindowing routine takes 16 of each product's bits,
rounded down, so each sum of 24 products came out about 12 units
low. Through the filter bank that put a dc offset of about 12 LSB
(of 16 bits) on the output and, from the high band, a tone at half
the sample rate; together they were four fifths of the decoder's
error on these targets.
Start the sums 12 up. Also correct a comment in both ARM files.
Checked with perfsim (Sansa Clip+ build) against ffmpeg's decode of
seven files: the dc offset goes from 11 to 12 LSB down to under
1 LSB, and the error from about -66 dBFS to -69 to -72 dBFS, the
same as the ARMv4 routine gives. Decoding is 0.7% slower (26.73 to
26.91 MHz for atrac3_lp2_132.oma). The ARMv4 output is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
applyVariableGain() applied the constant gain before a gain point
in a do-while loop, so at least eight samples of it, even when
there are none: when the point is at the start of the block or
follows straight after the previous one. Every later gain point in
the block was then eight samples late. This came in with the loop
unrolling of 51a8be1a0f.
Use a while loop, as ffmpeg does.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode.
Over the whole of atrac3_lp2_132.oma the SNR goes from 45 to 51 dB
and the worst frame from 14 to 21 dB; the right channel of a joint
stereo RM file goes from 37 to 54 dB. What is left is a noise floor
near -72 dBFS in every frame, from the 2 fractional bits the
decoder works with.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Only the first and last 128 points of the 512 point window were
applied, and the 256 between were taken to be one. They are not:
the window rises to 1.207 there. Every block was therefore up to
1.6 dB low over half its length, which left the output 0.7 dB low
on average and an error about 23 dB below the signal at all
frequencies.
Add the missing half of the table and apply it.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode,
with the two previous fixes applied before and after:
level SNR
five plain stereo RM files -0.7 dB 23 dB -> 0.0 dB 54-57 dB
joint stereo RM file -0.7 dB 23 dB -> 0.0 dB 37-45 dB
joint stereo OMA file -0.7 dB 23 dB -> 0.0 dB 45 dB
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The spectra of the odd QMF bands are stored reversed. 8722c6f2bb
moved that reversal out of IMLT() and into the decoding of the
coefficients, but did it within each scale factor band, one place
off, and not at all for tonal components. The odd bands, which is
5.5 to 11 kHz and above 16.5 kHz at 44.1 kHz, have been decoded
wrongly since: against ffmpeg's decode that range was uncorrelated.
Reverse the whole band in IMLT() again, as ffmpeg does. The old
commit measured the saving it is giving up as 0.11 MHz on PP5024.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode
of an OMA file and six RM files: the coherence of the 5.7 to 10.7
kHz range goes from under 0.01 to 0.99, the same as other bands.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Since 2e314093c8 the ATRAC3 decoder takes its frame size from
id3->bytesperframe, which the RM metadata parser does not set. With
a frame size of zero, a file in plain stereo mode had its second
channel decoded from the first channel's data: both outputs were
the left channel.
Set it from the RM sub packet size before the decoder is opened.
Checked with perfsim (Sansa e200v1 build) on six ATRAC3 .rm files:
the two channels were identical sample for sample before, and each
now follows its own channel of ffmpeg's decode.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The comment packet is not used by the codec, but skipping it grew
the stream's packet buffer to the packet's full size. Heap use rose
with the size of the tags: a file with 300 KB of embedded album art
needed 617 KB of heap, and one with 940 KB of art could not be
opened even with a 992 KB codec buffer.
Skip the packet without storing it: forget the part already
buffered and let ogg_stream_pagein() drop the rest, as it does for
a packet whose start was lost.
Seeking to the start of a file ran the headers through the stream
again, which buffered the whole comment packet by the normal path.
Start the seek search at the first audio page instead, as libvorbis
does, and handle a target inside that page.
With both, heap use no longer depends on the tags. Large tags cost
at most about 90 KB over an untagged file, from the two Ogg buffers
growing to hold one full page each.
Checked with perfsim on the Sansa Clip+ model, not on a device,
with the heap limited to what a 512 KB codec buffer leaves:
files with up to 1.5 MB of art or 240 KB of text decode to the
same PCM hashes as the untagged audio, and 15 seeks per file,
including to 0, 30 and 150 ms, give the same positions and PCM
as the unmodified codec on ten files. Chained streams were not
tested.
Change-Id: Id27ebf98b436d813cfa3179951eee5bd7b2c0cd2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Where the LSP curve peaks, the value its -1/4 power is taken of is
far below one step of 16.16 and was rounded to zero. pow_m1_4() of
zero is enormous, the block's levels are taken relative to that
maximum, and so the whole block came out about 46 dB down. In
StutteringFile.wma (22 kHz mono, 20 kbps) this muted a block here
and there, which is heard as a stutter.
Pass the value to pow_m1_4() with all of its 38 fractional bits.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode
of that file, in blocks of 46 ms: 11 of 841 were off by more than
1 dB, the worst by 10.6 dB; none are now, the worst by 0.4 dB.
Four other files with LSP exponents are unchanged or closer, and
files without them are byte-identical.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tremor is built here with CHANNELS 2: synthesis produces no PCM
for a stream with more channels, so a 5.1 file "played" as nothing
at all. It was still set up in full first, which took 416 to 880 KB
of the codec heap for the 5.1 and 7.1 files tried, and wrote one
entry per channel into arrays sized for two.
Refuse such a stream when its identification header is read. The
codec then returns an error at once, with under 8 KB of heap used.
Checked with perfsim on the Sansa Clip+ model, not on a device:
six 5.1 and 7.1 files are refused, and 13 stereo and mono files
give the same PCM hashes as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Very low bitrate WMA files send their spectral envelope as LSP
coefficients. The curve computed from them squares two products
whose value can reach about a million, in 16.16 fixed point. Where
the envelope should be lowest the squares overflowed, and those
points came out near the curve's maximum instead: bands at the top
of the spectrum up to 17 dB too loud, and everything else a few dB
low, as levels are taken relative to the maximum.
Sum the squares in 64 bits, and let pow_m1_4() take a value above
the 16.16 range.
Checked with perfsim (Sansa e200v1 build). The curve is within
0.2 dB of ffmpeg's formula at all 57600 points of 80 blocks traced
from two files. Against ffmpeg's decode, as the mean difference in
band levels over the whole file, with the noise coding fix applied
before and after:
StutteringFile (22 kHz mono, 20 kbps) 1.6 dB -> 0.3 dB
shoujicomcn2117164 (same) 0.2 dB -> 0.1 dB
terra8k (11 kHz mono, 8 kbps) 0.3 dB -> 0.0 dB
Only the five files that use LSP exponents change; the other 23
are byte-identical.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Low bitrate WMA files code their upper bands as noise at a given
power. Those bands came out 12 to 25 dB too quiet, which dulled
the treble. Two things were wrong.
The exponent pointer was moved too far in blocks whose exponents
have another resolution. This is ffmpeg's r20756 (f78501b264,
"Fix apparent 10l typos introduced in r8627"); r8627 was merged
here in 2f1da8d24a but the fix for it never was.
The fixed point gain of a band lost nearly all its precision: it
was zero for three bands in four in the file traced, and the sum
for a band's power overflowed in a quarter of them. Compute the
gain in 64 bits, once a band. The factor for the noise below the
first coded coefficient (WMA v1 only) was 16 bits too large and is
corrected by the same reasoning; no file here exercises it.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode,
as the mean difference in band levels over the whole file:
01 - Jane Austen (44.1 kHz, 32 kbps) 5.5 dB -> 0.0 dB
07.Devil.In.My.Mind (same) 6.2 dB -> 0.0 dB
beyonthepain907z (same) 6.8 dB -> 0.0 dB
moshimoashitaga (same) 2.8 dB -> 0.0 dB
Four more noise coded files, already within 0.3 dB, now match too.
The output of the 15 files without noise coding is byte-identical.
Files that use LSP exponents still differ in their top bands; that
is a separate fault in the LSP curve.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A payload normally holds one codec packet of blockalign bytes, but
some files put several in each. Only the first was decoded, so such
a file played one packet in every 8 or 15, as a few seconds of
broken sound, and then ended.
Step through the payload in blockalign sized packets.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode:
doors-test.wma (WMA v2, 11 kHz mono), 15 packets a payload:
80896 of 1205760 samples before, all of them after, 86 dB SNR
test.wma (WMA v1, 44.1 kHz stereo), 8 packets a payload:
1294336 of 10346496 samples before, all after, 114 dB SNR
The output of 26 other WMA files is byte-identical before and after.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A WMA Professional stream of 16 bits per sample played about 48 dB
too quiet and with a noise floor near -70 dBFS.
ffmpeg scales the transform's output by the stream's sample size.
That was dropped when the decoder was converted to fixed point
(d884af2b99, 16284ae8ae), and the output is passed to the DSP as if
every stream had 24 bits. A stream of fewer bits has a lower
quantization step to match, so it came out low by the difference,
and it used the bottom of the integer quantization table, where the
factors have only a few significant bits.
Decode a 16 or 20 bit stream at the level of a 24 bit stream: use
that stream's quantization step, and scale each band's factor by
the ratio that is left. 24 bit streams are not affected.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode
of a 16 bit, 192 kbps stereo file, the only such file to hand:
level SNR vs ffmpeg noise
before -48 dB 37 dB -70 dBFS
after correct 85 dB -117 dBFS
A 24 bit file's output is byte-identical before and after. The
20 bit case follows the same rule but is untested.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
inverse_channel_transform() ran its general N-channel matrix loop
for every sample of a stereo stream, where the matrix is exactly
+-1.0. That loop was a quarter to a third of the whole decode, and
GCC 9.5.0 compiles it worse than 4.9.4 did, which made the codec
6-9% slower on ARM7TDMI after the toolchain update.
Handle a group of two channels separately: add and subtract when
the matrix is +-1.0, and a plain four multiply loop otherwise. More
than two channels still use the general loop.
This reverses the regression from the GCC 9.5.0 update and goes
well past it. The loop the newer compiler handled badly is no longer
used for stereo, so the two compilers now give the same speed to
within 1%, about 25% faster than the codec was with GCC 4.9.4
(estimated with perfsim, e200v1, wmapro_141k: 25.81 MHz with 4.9.4
before this change, 19.5 MHz with either compiler after it).
Output is bit-identical: whole-file PCM hashes match before and
after for five stereo files at 55-271 kbps, built with GCC 9.5.0
and with 4.9.4, and also with the multiply path forced on.
Measured with test_codec, wmapro_141k.wma, MHz for real time:
Sansa e200v1 27.99 -> 19.70
Sansa Clip+ 21.78 -> 15.80
Estimated with perfsim for the other files (e200v1 / Clip+):
wmapro_55k 25.21 -> 17.06 / 20.02 -> 13.71
wmapro_80k 26.17 -> 18.01 / 20.75 -> 14.44
wmapro_173k 28.52 -> 20.29 / 22.55 -> 16.25
wmapro_271k 30.85 -> 22.52 / 24.34 -> 18.04
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replaces flac_lpc_32_c, the wide predictor path every 24-bit stream
takes, and 16-bit streams encoded at high coefficient precision. The
coefficients are invariant for the whole call, but the C loop reloaded
all of them for every output sample and spilled its loop bound to the
stack on top of that. Orders 1-8 now keep every coefficient in a
register; orders 10 and 12, which is what -8 emits, get their own
unrolled loops instead of the generic chunked one. An ARMv5E kernel
using the packed 16-bit multiplies is included behind
FLAC_LPC32_NARROW_ASM, off by default: it needs bps <= 16, and only
about half the subframes of such a stream can use it.
Measured with test_codec on 24-bit/96kHz streams. At predictor order 12:
36.24 MHz on e200 (ARMv4) against 62.29 MHz before. The same file
improves from 37.65 MHz to 27.3 MHz on Clip+ (ARMv5).
Bit-exact against flac -d over four complete streams on both targets,
with and without the assembly, and the kernels are checked against an
int64_t reference across every order 1-32, qlevel 0-15 and coefficient
precision 1-15.
The rarely-executed orders stay in DRAM: the codec's IRAM window on
PP502x is nearly full and demoting them measured no cost, leaving 96
bytes free. FLAC_LPC32_NO_IRAM demotes the whole filter.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Icd8bd298973b9f5c26c103af8adf47b03489d69f
Of the two divisions a sample left in the entropy decoder, the first
divides the range by the Rice pivot, which is small: below 1024 for
99.9% of the samples of a 16-bit test file and never as much as 4096.
Keep a table of reciprocals, filled in as divisors turn up, and
divide with a 32x32->64 multiply and one correction. Larger divisors
go to the division routine as before. The other division, by help,
has no such pattern.
This is for ARMv5 and later without a hardware divide. ARMv4 has
its own divider with a reciprocal table already, and a long multiply
is slow there; its codec is unchanged.
The table is 16 KB of bss. With r = (2^32 - 1) / n the estimate is
the quotient or one less for any 32-bit numerator, checked against
true division for every n below 4096.
MHz for real time in perfsim's model of the Clip+, -c1000, -c2000
and -c3000: 32.5, 48.4, 76.4 to 30.1, 46.0, 74.0. A Clip+ measures
30.63 at -c1000, from 32.85, with the same checksum, 1008ffab.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I555ff83d7e19919848c8fb7b8d12ecfa671a976b
After the division routines the entropy decoder was the largest part
of Monkey's Audio at the fast levels, with two things to gain.
All of its state, the range coder's and the Rice parameters, was in
statics, so each use was a load and each update a store: about half
of the function's time on ARM7TDMI, where a load is 3 cycles. Copy
it to locals for the length of a block and write it back after. The
functions that take a pointer to it have to be inlined for that to
work, and on ARMv4 gcc left the per-sample one out of line, so force
them.
The symbol was found by dividing low by help and searching the count
table for the quotient. counts[n] <= low / help is the same as
counts[n] * help <= low, so search with the multiply instead: the
first symbols are by far the likeliest. That leaves two divisions a
sample from three.
MHz for real time in perfsim's models, -c1000, -c2000 and -c3000:
Clip+ (ARMv5) 37.9, 53.8, 81.8 to 32.5, 48.4, 76.4
e200 (ARMv4) 45.3, 68.6, 111.4 to 35.2, 58.5, 101.3
A Clip+ measures 32.85 at -c1000 (61.4 before this and the division
change), with test_codec's checksum, 1008ffab, the same as the old
code's. The standalone decoder gives identical output for all four
levels tested. The path for files older than 3.98 has the same
change and was not tested, for want of a file.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Idf9002a5b5f3a423f44f70443936ea5ac427cd12
lib/arm_support/support-arm.S was written to replace libgcc's
division for ARM, and bff5a35c3c (FS#10943, 2010) added it to the
core, the plugin library and the codec library alike. When
1501df045f (2013) replaced EXTRA_LIBS with explicit lists, plugins
kept it and codecs did not, and they have taken their division from
libgcc since.
With the gcc 4.4 toolchain that cost little: its libgcc had a
routine that used clz. With gcc 9.5 libgcc has no soft-float ARMv5
variant, so ARMv5 targets get the ARMv4 routine, a shift and
subtract loop of about 130 cycles a division.
Monkey's Audio divides two or three times a sample in its range
decoder and, without codec IRAM, does it in C. On a Clip+ it is
11% to 21% slower than 3.14 was, with over half of -c1000's decode
in __udivsi3. MP2 is 4% to 7% slower.
Put libarm_support back, ahead of libgcc. In perfsim's model of
the Clip+ a division falls to 44 cycles, Monkey's Audio by 38%, 30%
and 22% at -c1000, -c2000 and -c3000 (60.9 to 37.9 MHz at -c1000),
and MP2 by 3% to 7%; nothing else moves by more than 0.6%. A Clip+
decodes -c1000 with the same checksum as before.
A division by zero in a codec goes to __div0 again, and so to the
firmware's handler, as it did with support-arm.S and with the old
libgcc (3.14's ape.codec calls it). The gcc 9.5 libgcc returns
from its own stub instead.
Change-Id: I6a9a79ce8e85bca69870349a5c0823f392a578b6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The frame decoder reads from one flat buffer and cannot refill it, but
request_buffer() only guarantees 32KiB of contiguous data (the buffering
guard area). Frames that can be larger than that, such as high
resolution or poorly compressible streams (FLAC decoder testbench file
31), could be handed to the decoder truncated, which read past the end
of the data and lost sync.
When a request returns less than the largest frame the stream can
contain (STREAMINFO max framesize, or a bound from block size, channels
and bit depth) and it is not the end of the file, copy the frame into a
private static buffer and decode from that. Streams whose frames always
fit never touch the buffer and pay one comparison per frame. The buffer
is 64KiB, or sized for the 4608 sample blocks of memory limited targets,
and is left out entirely when MEMORYSIZE is 2MB or less.
Also fail with a codec error, instead of advancing past the data, if a
decoded frame consumed more bytes than were provided.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: Ib4cf85ad513f8e096f73d13b4374d072e1745cb2
At a CELT to hybrid mode switch, opus_decode_frame() runs CELT's pitch
PLC nested in the new frame, and celt_decode_lost() held a copy of up to
2 KB of excitation for celt_fir(). On stackOverflow.opus that overran
the 9 KB codec stack on native targets such as the e200v2. Filtering in
place from the last sample down needs no copy; output is bit-identical.
Worst case over all 242 mode switches of stackOverflow.opus, measured
under qemu: 9140 -> 7676 bytes below opus_decode(), against 8912 left
for it on the e200v2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Iaa9f57182ad86ab76c81c55bea72aa517cc18230
The folded Rice value was unfolded with a signed shift, which is wrong
once it reaches 2^31, and the unary length limit was (INT_MAX >> k) + 2,
about half of what a 32-bit residual can need. Streams with very large
residuals (FLAC decoder testbench file 63) were misparsed, overran the
frame and lost sync at the next one. Unfold as unsigned and derive the
limit from UINT_MAX, clamped to INT_MAX. The fast path is unchanged.
The existing 0x80000000 error check now also works as intended, since
the Golomb reader's error value maps to it. The FLAC spec forbids a
residual of -2^31.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: I0bdce9db69e1b30565ecc23126923ba401b7deca
On ARMv5E and later the fixed-phase FIR cycles at 8, 12 and 16 kHz load
two samples a word and take two coefficients a word from a literal
pool; smla<x><y> picks the halves, so an odd-aligned window costs
nothing. Outputs are paired as two interleaved accumulator chains. The
FIR buffer is now word aligned. Generated by
silk/arm/gen_resampler_armv5e.py. Bit-exact; OPUS_ARM_NO_SILK_ASM
disables it, and config.h sets that on M-profile cores.
Measured on the Clip+, silk_5.opus (WB SILK):
20.93 -> 15.84 MHz, -24.3%.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Ibec71fd52542f768f22e2970c9f8c45118c708b9
On ARMv4 the fixed-phase FIR cycles at 8, 12 and 16 kHz run as adds of
shifted samples rather than multiplies: every coefficient is a constant,
and a shifted add is one cycle where mla plus loading the coefficient is
five or six. Each sample is loaded once per cycle and added into the
two or three outputs it feeds, sharing partial products such as 31x
between them, about 27 adds per output. The kernels are generated by
silk/arm/gen_resampler_armv4.py. Bit-exact; OPUS_ARM_NO_SILK_ASM
disables them.
Measured on the e200v1, silk_5.opus (WB SILK), with the previous commit:
36.31 -> 28.85 MHz, -20.5%.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Ic1fd5777e91d182248f43954e4b6ae977309dc5d
The 8, 12 and 16 kHz to 48 kHz steps visit only two or three FIR phases
in a fixed cycle, so each set of input samples is read once and reused
across outputs. Bit-exact; OPUS_NO_SILK_FIXED_PHASE disables it.
Measured with silk_5.opus (WB SILK):
e200v1: 36.31 -> 33.38 MHz, -8.1%
Clip+: 21.66 -> 20.93 MHz, -3.4%
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Icb0f8ded62739dfee1f574e4888a54c778aa3d53
A total sample count of 0 means unknown. It made the track length 0 and the
bitrate estimate in flac_init() divided by it, crashing the codec (FLAC
decoder testbench file 45). Report a bitrate of 0 in that case.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
An escape code with a raw bit width of 0 means every residual in the
partition is zero. get_sbits(&gb, 0) shifts by 32, which is undefined and
returned stale cache bits instead of 0, so such streams decoded to garbage
(FLAC decoder testbench file 64).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
print_mp3entry() dereferences mb_track_id, but get_metadata() does not set
every field of the uninitialized stack struct. For FLAC files this left
garbage in the pointer and warble segfaulted in strlen about a third of the
time, before decoding started. Clear the struct first.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Upstream's OPUS_ARM_INLINE_ASM means any ARM with inline assembly, with
OPUS_ARM_INLINE_EDSP layered on top for ARMv5E. config.h instead defines
exactly one of them per core, so the names read as broader than they are,
and code added ARM_ARCH tests beside them to pin the scope down. Renamed
to OPUS_ARM_ASM_ARMV4_ONLY and OPUS_ARM_ASM_ARMV5E_AND_LATER throughout
celt and silk, upstream files included; README.rockbox records it for the
next libopus sync.
No code change: opus.elf disassembly and section sizes are identical
before and after on ARMv4 (e200v1), ARMv5E (Clip+) and ARMv6 (iPod Nano
4G).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I89932d348dedd76748a1bc7b3f2e7de1c5be49c8
Four `ARM_ARCH == 5` checks -- in SOURCES and celt/arm/{bands_arm,
comb_filter_arm,vq_arm}.h -- excluded ARMv6 from every ARMv5E kernel:
denorm_band, haar1, comb_filter_const, celt_sat, deemphasis_stereo_simple,
exp_rotation1 and normres_scale all silently fell back to plain C on
ARM1136/ARM1176, since ARMv6 is a strict superset of the EDSP instructions
those kernels use. Widened to ARM_ARCH >= 5, matching config.h's own
OPUS_ARM_INLINE_EDSP gate, which was already ARM_ARCH > 4.
These four can't be fixed at the commits that introduced them: those
commits are already merged into master under different SHAs. A fifth
instance of the same bug, in celt/arm/mdct_armv5e.h, was fixed at its
origin commit instead, since that one is still open for review.
Verified on both native ARMv6 targets (iPod Nano 4G, ARM1176JZ-S; Gigabeat
S, ARM1136JF-S) and the hosted Samsung YP-R0 (ARM1176JZ-S, cross toolchain
built for the occasion): all three now link and call all 16 ARMv5E
kernels, where they linked and called zero before this fix. Decoded PCM
bit-identical to the ARMv5E build under qemu.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: I7e6f6b421b9c0203812fb45d2258c7eb81738d98
53% of the dimensions a decode walks hold no pulses, and that arm leaves
_k alone, so the two CELT_PVQ_U_ROW pointers stay valid. U is
non-decreasing in _k, so p <= _i < q is the single unsigned test
(_i-p) < (q-p). Bit-exact over 8.2M samples.
Modelled: -0.40% ARMv4, -0.78% ARMv5E; the function -4.8% and -7.1%.
Measured: e200v1 39.19 -> 38.96 MHz, Clip+ 28.13 -> 28.08 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id50be222e101aeabd9a1f4cf5380d3244da610ff
At stride 1 the rotation chain writes X[i+stride] and reads it straight
back, so one load an iteration is redundant and one store is dead. The
kernel carries that value, narrowed, since mul reads all 32 bits where
smulbb does not. Bit-exact over 8.2M samples.
Modelled: -0.87% ARMv4, the function -10.0%.
Measured: e200v1 39.59 -> 39.19 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Idbd1a088bc805fecfb9ee54fc372ad16a7d91606
34 functions and the ARMv4 kernels, chosen by a greedy fill of the free
codec IRAM window ranked by cycles per byte. The mixed-radix
butterflies give their IRAM back, being unreachable under the prime
factor transform.
The cycle model does not see this at all: it models no instruction cache.
Measured: e200v1 48.96 -> 42.60 MHz; reclaiming the mixed-radix IRAM was
a further -0.52%.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I9e22025cba99521750a3b664cd0d59357fcd0810
Every 48 kHz CELT length is 15 times a power of two and the factors are
co-prime, so the inter-stage twiddles -- 73% of the FFT multiplies at
N=480 -- vanish. New pfa_fft15 and mdct_postrot_pfa kernels on both
cores; accuracy also improves 0.3 to 0.4 dB against opusdec.
Modelled: -6.17% ARMv4, -1.82% ARMv5E.
Measured: e200v1 42.33 -> 40.10 MHz, Clip+ 29.30 -> 28.24 MHz.
Rescheduling these kernels, folded in here, measured a further -0.75% on
e200v1 and -0.39% on Clip+.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ifa4a30d1045e905522828958949954615faa440d
One ldr fetches two celt_norm coefficients and smlabb/smlatt take the
halves apart; scalar fallback when the three pointers disagree on
alignment.
Modelled: -0.88% ARMv5E.
Measured: Clip+ 28.13 -> 28.08 MHz, the same build with and without it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I79719132ccecfb7664616bca019e3856ee3479da
Rockbox lists no encoder, so CELT_DECODE_ONLY folds away the thirteen
encoder branches in celt/bands.c and stops gcc keeping their values live
across quant_partition's recursive calls. 2,976 bytes smaller.
Modelled: -0.47% ARMv4, -0.69% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ie96a0a63e035ab0cdcdbe9eeeb6691f24fecfa41
Its reachable input set is 2,985 (qn,i) pairs for any stream ever, so it
tabulates exactly in 6,484 bytes. Verified exhaustively against the
compiled functions. Only built for PP5022/PP5024, where the tables fit
the 80 KB IRAM window; elsewhere it measured no gain.
Measured, e200v1: 39.18 -> 38.96 MHz with the tables in IRAM, 39.07 with
them in DRAM. Clip+: 28.05 -> 28.08 MHz, so not built there.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I38417793dc337ec6cb61139cc919d23ac09b04dd
This re-enables upstream ASM optimizations for Cortex-M. Only the
our downstream improvements are disabled, as they do not assemble in
thumb2 mode.
Change-Id: Icca1b3dbf04786c7714fb4eef92ad66aa55132f3
* Only use new ARMv5e optimizations on classic (non-M) profile
* fix inconsistent ARM_ARCH >= 5 vs == 5
Fixes red in 0c4345475a and ae223933bf
Change-Id: Ieb688679d2d698a19870b1a63b4151a8593e04e0
radix-3, 4 and 5. The gain is bookkeeping: twiddles addressed by
displacement from one base register, post-indexed stores, and C_MUL's
Q15 doubling folded into the add that consumes it.
Modelled: -4.72% ARMv5E, opus_fft_impl -25.3%.
Measured: Clip+ 30.89 -> 29.33 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I173667ebb13308f83bd6fcb4e876f5a56c698cd3
Pre-rotation, post-rotation and mirror, with a packed-twiddle complex
multiply throughout.
Modelled: -8.0% ARMv5E against the C loops.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I3fcaf2343e04e3f3f4462fc442f8fac623729631
deemphasis_stereo_simple on both cores, with the filter state kept
unshifted and shifted inside the add that consumes it.
Modelled: -0.53% ARMv4, -0.96% ARMv5E.
Measured with the three preceding commits: e200v1 49.42 -> 48.96 MHz,
Clip+ 32.83 -> 30.89 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id09297974f3b66a5a193b524d68fd8844f71cc02
ARMv4 has no CLZ, so all 8,680 ilog2 calls went through libgcc's
__clzsi2. Fifteen branchless instructions replace it.
Modelled: -0.91% ARMv4; ARMv5E unaffected, it already emits CLZ.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ib84b97acb687da4101e02bad1b0e8bbd256d32df
comb_filter_const on both cores, and celt_synthesis's SIG_SAT clamp
four samples at a time through one ldm and one stm.
Modelled for the saturation loop: -1.05% ARMv4, -1.42% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I381afab451eb29a91513ce6ab29c1b2d565b7ee9
denormalise_bands on both cores; exp_rotation1, haar1 and the
normalise_residual scaling loop on ARMv5E. Bit-exact.
Modelled: -0.50% ARMv4, -2.91% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: If40f5dbe7cb9e7c12aa0a5b4ac9e73433a850b6e