inverse_channel_transform() ran its general N-channel matrix loop
for every sample of a stereo stream, where the matrix is exactly
+-1.0. That loop was a quarter to a third of the whole decode, and
GCC 9.5.0 compiles it worse than 4.9.4 did, which made the codec
6-9% slower on ARM7TDMI after the toolchain update.
Handle a group of two channels separately: add and subtract when
the matrix is +-1.0, and a plain four multiply loop otherwise. More
than two channels still use the general loop.
This reverses the regression from the GCC 9.5.0 update and goes
well past it. The loop the newer compiler handled badly is no longer
used for stereo, so the two compilers now give the same speed to
within 1%, about 25% faster than the codec was with GCC 4.9.4
(estimated with perfsim, e200v1, wmapro_141k: 25.81 MHz with 4.9.4
before this change, 19.5 MHz with either compiler after it).
Output is bit-identical: whole-file PCM hashes match before and
after for five stereo files at 55-271 kbps, built with GCC 9.5.0
and with 4.9.4, and also with the multiply path forced on.
Measured with test_codec, wmapro_141k.wma, MHz for real time:
Sansa e200v1 27.99 -> 19.70
Sansa Clip+ 21.78 -> 15.80
Estimated with perfsim for the other files (e200v1 / Clip+):
wmapro_55k 25.21 -> 17.06 / 20.02 -> 13.71
wmapro_80k 26.17 -> 18.01 / 20.75 -> 14.44
wmapro_173k 28.52 -> 20.29 / 22.55 -> 16.25
wmapro_271k 30.85 -> 22.52 / 24.34 -> 18.04
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replaces flac_lpc_32_c, the wide predictor path every 24-bit stream
takes, and 16-bit streams encoded at high coefficient precision. The
coefficients are invariant for the whole call, but the C loop reloaded
all of them for every output sample and spilled its loop bound to the
stack on top of that. Orders 1-8 now keep every coefficient in a
register; orders 10 and 12, which is what -8 emits, get their own
unrolled loops instead of the generic chunked one. An ARMv5E kernel
using the packed 16-bit multiplies is included behind
FLAC_LPC32_NARROW_ASM, off by default: it needs bps <= 16, and only
about half the subframes of such a stream can use it.
Measured with test_codec on 24-bit/96kHz streams. At predictor order 12:
36.24 MHz on e200 (ARMv4) against 62.29 MHz before. The same file
improves from 37.65 MHz to 27.3 MHz on Clip+ (ARMv5).
Bit-exact against flac -d over four complete streams on both targets,
with and without the assembly, and the kernels are checked against an
int64_t reference across every order 1-32, qlevel 0-15 and coefficient
precision 1-15.
The rarely-executed orders stay in DRAM: the codec's IRAM window on
PP502x is nearly full and demoting them measured no cost, leaving 96
bytes free. FLAC_LPC32_NO_IRAM demotes the whole filter.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Icd8bd298973b9f5c26c103af8adf47b03489d69f
Of the two divisions a sample left in the entropy decoder, the first
divides the range by the Rice pivot, which is small: below 1024 for
99.9% of the samples of a 16-bit test file and never as much as 4096.
Keep a table of reciprocals, filled in as divisors turn up, and
divide with a 32x32->64 multiply and one correction. Larger divisors
go to the division routine as before. The other division, by help,
has no such pattern.
This is for ARMv5 and later without a hardware divide. ARMv4 has
its own divider with a reciprocal table already, and a long multiply
is slow there; its codec is unchanged.
The table is 16 KB of bss. With r = (2^32 - 1) / n the estimate is
the quotient or one less for any 32-bit numerator, checked against
true division for every n below 4096.
MHz for real time in perfsim's model of the Clip+, -c1000, -c2000
and -c3000: 32.5, 48.4, 76.4 to 30.1, 46.0, 74.0. A Clip+ measures
30.63 at -c1000, from 32.85, with the same checksum, 1008ffab.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I555ff83d7e19919848c8fb7b8d12ecfa671a976b
After the division routines the entropy decoder was the largest part
of Monkey's Audio at the fast levels, with two things to gain.
All of its state, the range coder's and the Rice parameters, was in
statics, so each use was a load and each update a store: about half
of the function's time on ARM7TDMI, where a load is 3 cycles. Copy
it to locals for the length of a block and write it back after. The
functions that take a pointer to it have to be inlined for that to
work, and on ARMv4 gcc left the per-sample one out of line, so force
them.
The symbol was found by dividing low by help and searching the count
table for the quotient. counts[n] <= low / help is the same as
counts[n] * help <= low, so search with the multiply instead: the
first symbols are by far the likeliest. That leaves two divisions a
sample from three.
MHz for real time in perfsim's models, -c1000, -c2000 and -c3000:
Clip+ (ARMv5) 37.9, 53.8, 81.8 to 32.5, 48.4, 76.4
e200 (ARMv4) 45.3, 68.6, 111.4 to 35.2, 58.5, 101.3
A Clip+ measures 32.85 at -c1000 (61.4 before this and the division
change), with test_codec's checksum, 1008ffab, the same as the old
code's. The standalone decoder gives identical output for all four
levels tested. The path for files older than 3.98 has the same
change and was not tested, for want of a file.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Idf9002a5b5f3a423f44f70443936ea5ac427cd12
lib/arm_support/support-arm.S was written to replace libgcc's
division for ARM, and bff5a35c3c (FS#10943, 2010) added it to the
core, the plugin library and the codec library alike. When
1501df045f (2013) replaced EXTRA_LIBS with explicit lists, plugins
kept it and codecs did not, and they have taken their division from
libgcc since.
With the gcc 4.4 toolchain that cost little: its libgcc had a
routine that used clz. With gcc 9.5 libgcc has no soft-float ARMv5
variant, so ARMv5 targets get the ARMv4 routine, a shift and
subtract loop of about 130 cycles a division.
Monkey's Audio divides two or three times a sample in its range
decoder and, without codec IRAM, does it in C. On a Clip+ it is
11% to 21% slower than 3.14 was, with over half of -c1000's decode
in __udivsi3. MP2 is 4% to 7% slower.
Put libarm_support back, ahead of libgcc. In perfsim's model of
the Clip+ a division falls to 44 cycles, Monkey's Audio by 38%, 30%
and 22% at -c1000, -c2000 and -c3000 (60.9 to 37.9 MHz at -c1000),
and MP2 by 3% to 7%; nothing else moves by more than 0.6%. A Clip+
decodes -c1000 with the same checksum as before.
A division by zero in a codec goes to __div0 again, and so to the
firmware's handler, as it did with support-arm.S and with the old
libgcc (3.14's ape.codec calls it). The gcc 9.5 libgcc returns
from its own stub instead.
Change-Id: I6a9a79ce8e85bca69870349a5c0823f392a578b6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The frame decoder reads from one flat buffer and cannot refill it, but
request_buffer() only guarantees 32KiB of contiguous data (the buffering
guard area). Frames that can be larger than that, such as high
resolution or poorly compressible streams (FLAC decoder testbench file
31), could be handed to the decoder truncated, which read past the end
of the data and lost sync.
When a request returns less than the largest frame the stream can
contain (STREAMINFO max framesize, or a bound from block size, channels
and bit depth) and it is not the end of the file, copy the frame into a
private static buffer and decode from that. Streams whose frames always
fit never touch the buffer and pay one comparison per frame. The buffer
is 64KiB, or sized for the 4608 sample blocks of memory limited targets,
and is left out entirely when MEMORYSIZE is 2MB or less.
Also fail with a codec error, instead of advancing past the data, if a
decoded frame consumed more bytes than were provided.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: Ib4cf85ad513f8e096f73d13b4374d072e1745cb2
At a CELT to hybrid mode switch, opus_decode_frame() runs CELT's pitch
PLC nested in the new frame, and celt_decode_lost() held a copy of up to
2 KB of excitation for celt_fir(). On stackOverflow.opus that overran
the 9 KB codec stack on native targets such as the e200v2. Filtering in
place from the last sample down needs no copy; output is bit-identical.
Worst case over all 242 mode switches of stackOverflow.opus, measured
under qemu: 9140 -> 7676 bytes below opus_decode(), against 8912 left
for it on the e200v2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Iaa9f57182ad86ab76c81c55bea72aa517cc18230
The folded Rice value was unfolded with a signed shift, which is wrong
once it reaches 2^31, and the unary length limit was (INT_MAX >> k) + 2,
about half of what a 32-bit residual can need. Streams with very large
residuals (FLAC decoder testbench file 63) were misparsed, overran the
frame and lost sync at the next one. Unfold as unsigned and derive the
limit from UINT_MAX, clamped to INT_MAX. The fast path is unchanged.
The existing 0x80000000 error check now also works as intended, since
the Golomb reader's error value maps to it. The FLAC spec forbids a
residual of -2^31.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: I0bdce9db69e1b30565ecc23126923ba401b7deca
On ARMv5E and later the fixed-phase FIR cycles at 8, 12 and 16 kHz load
two samples a word and take two coefficients a word from a literal
pool; smla<x><y> picks the halves, so an odd-aligned window costs
nothing. Outputs are paired as two interleaved accumulator chains. The
FIR buffer is now word aligned. Generated by
silk/arm/gen_resampler_armv5e.py. Bit-exact; OPUS_ARM_NO_SILK_ASM
disables it, and config.h sets that on M-profile cores.
Measured on the Clip+, silk_5.opus (WB SILK):
20.93 -> 15.84 MHz, -24.3%.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Ibec71fd52542f768f22e2970c9f8c45118c708b9
On ARMv4 the fixed-phase FIR cycles at 8, 12 and 16 kHz run as adds of
shifted samples rather than multiplies: every coefficient is a constant,
and a shifted add is one cycle where mla plus loading the coefficient is
five or six. Each sample is loaded once per cycle and added into the
two or three outputs it feeds, sharing partial products such as 31x
between them, about 27 adds per output. The kernels are generated by
silk/arm/gen_resampler_armv4.py. Bit-exact; OPUS_ARM_NO_SILK_ASM
disables them.
Measured on the e200v1, silk_5.opus (WB SILK), with the previous commit:
36.31 -> 28.85 MHz, -20.5%.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Ic1fd5777e91d182248f43954e4b6ae977309dc5d
The 8, 12 and 16 kHz to 48 kHz steps visit only two or three FIR phases
in a fixed cycle, so each set of input samples is read once and reused
across outputs. Bit-exact; OPUS_NO_SILK_FIXED_PHASE disables it.
Measured with silk_5.opus (WB SILK):
e200v1: 36.31 -> 33.38 MHz, -8.1%
Clip+: 21.66 -> 20.93 MHz, -3.4%
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Icb0f8ded62739dfee1f574e4888a54c778aa3d53
A total sample count of 0 means unknown. It made the track length 0 and the
bitrate estimate in flac_init() divided by it, crashing the codec (FLAC
decoder testbench file 45). Report a bitrate of 0 in that case.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
An escape code with a raw bit width of 0 means every residual in the
partition is zero. get_sbits(&gb, 0) shifts by 32, which is undefined and
returned stale cache bits instead of 0, so such streams decoded to garbage
(FLAC decoder testbench file 64).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
print_mp3entry() dereferences mb_track_id, but get_metadata() does not set
every field of the uninitialized stack struct. For FLAC files this left
garbage in the pointer and warble segfaulted in strlen about a third of the
time, before decoding started. Clear the struct first.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Upstream's OPUS_ARM_INLINE_ASM means any ARM with inline assembly, with
OPUS_ARM_INLINE_EDSP layered on top for ARMv5E. config.h instead defines
exactly one of them per core, so the names read as broader than they are,
and code added ARM_ARCH tests beside them to pin the scope down. Renamed
to OPUS_ARM_ASM_ARMV4_ONLY and OPUS_ARM_ASM_ARMV5E_AND_LATER throughout
celt and silk, upstream files included; README.rockbox records it for the
next libopus sync.
No code change: opus.elf disassembly and section sizes are identical
before and after on ARMv4 (e200v1), ARMv5E (Clip+) and ARMv6 (iPod Nano
4G).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I89932d348dedd76748a1bc7b3f2e7de1c5be49c8
Four `ARM_ARCH == 5` checks -- in SOURCES and celt/arm/{bands_arm,
comb_filter_arm,vq_arm}.h -- excluded ARMv6 from every ARMv5E kernel:
denorm_band, haar1, comb_filter_const, celt_sat, deemphasis_stereo_simple,
exp_rotation1 and normres_scale all silently fell back to plain C on
ARM1136/ARM1176, since ARMv6 is a strict superset of the EDSP instructions
those kernels use. Widened to ARM_ARCH >= 5, matching config.h's own
OPUS_ARM_INLINE_EDSP gate, which was already ARM_ARCH > 4.
These four can't be fixed at the commits that introduced them: those
commits are already merged into master under different SHAs. A fifth
instance of the same bug, in celt/arm/mdct_armv5e.h, was fixed at its
origin commit instead, since that one is still open for review.
Verified on both native ARMv6 targets (iPod Nano 4G, ARM1176JZ-S; Gigabeat
S, ARM1136JF-S) and the hosted Samsung YP-R0 (ARM1176JZ-S, cross toolchain
built for the occasion): all three now link and call all 16 ARMv5E
kernels, where they linked and called zero before this fix. Decoded PCM
bit-identical to the ARMv5E build under qemu.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: I7e6f6b421b9c0203812fb45d2258c7eb81738d98
53% of the dimensions a decode walks hold no pulses, and that arm leaves
_k alone, so the two CELT_PVQ_U_ROW pointers stay valid. U is
non-decreasing in _k, so p <= _i < q is the single unsigned test
(_i-p) < (q-p). Bit-exact over 8.2M samples.
Modelled: -0.40% ARMv4, -0.78% ARMv5E; the function -4.8% and -7.1%.
Measured: e200v1 39.19 -> 38.96 MHz, Clip+ 28.13 -> 28.08 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id50be222e101aeabd9a1f4cf5380d3244da610ff
At stride 1 the rotation chain writes X[i+stride] and reads it straight
back, so one load an iteration is redundant and one store is dead. The
kernel carries that value, narrowed, since mul reads all 32 bits where
smulbb does not. Bit-exact over 8.2M samples.
Modelled: -0.87% ARMv4, the function -10.0%.
Measured: e200v1 39.59 -> 39.19 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Idbd1a088bc805fecfb9ee54fc372ad16a7d91606
34 functions and the ARMv4 kernels, chosen by a greedy fill of the free
codec IRAM window ranked by cycles per byte. The mixed-radix
butterflies give their IRAM back, being unreachable under the prime
factor transform.
The cycle model does not see this at all: it models no instruction cache.
Measured: e200v1 48.96 -> 42.60 MHz; reclaiming the mixed-radix IRAM was
a further -0.52%.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I9e22025cba99521750a3b664cd0d59357fcd0810
Every 48 kHz CELT length is 15 times a power of two and the factors are
co-prime, so the inter-stage twiddles -- 73% of the FFT multiplies at
N=480 -- vanish. New pfa_fft15 and mdct_postrot_pfa kernels on both
cores; accuracy also improves 0.3 to 0.4 dB against opusdec.
Modelled: -6.17% ARMv4, -1.82% ARMv5E.
Measured: e200v1 42.33 -> 40.10 MHz, Clip+ 29.30 -> 28.24 MHz.
Rescheduling these kernels, folded in here, measured a further -0.75% on
e200v1 and -0.39% on Clip+.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ifa4a30d1045e905522828958949954615faa440d
One ldr fetches two celt_norm coefficients and smlabb/smlatt take the
halves apart; scalar fallback when the three pointers disagree on
alignment.
Modelled: -0.88% ARMv5E.
Measured: Clip+ 28.13 -> 28.08 MHz, the same build with and without it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I79719132ccecfb7664616bca019e3856ee3479da
Rockbox lists no encoder, so CELT_DECODE_ONLY folds away the thirteen
encoder branches in celt/bands.c and stops gcc keeping their values live
across quant_partition's recursive calls. 2,976 bytes smaller.
Modelled: -0.47% ARMv4, -0.69% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ie96a0a63e035ab0cdcdbe9eeeb6691f24fecfa41
Its reachable input set is 2,985 (qn,i) pairs for any stream ever, so it
tabulates exactly in 6,484 bytes. Verified exhaustively against the
compiled functions. Only built for PP5022/PP5024, where the tables fit
the 80 KB IRAM window; elsewhere it measured no gain.
Measured, e200v1: 39.18 -> 38.96 MHz with the tables in IRAM, 39.07 with
them in DRAM. Clip+: 28.05 -> 28.08 MHz, so not built there.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I38417793dc337ec6cb61139cc919d23ac09b04dd
This re-enables upstream ASM optimizations for Cortex-M. Only the
our downstream improvements are disabled, as they do not assemble in
thumb2 mode.
Change-Id: Icca1b3dbf04786c7714fb4eef92ad66aa55132f3
* Only use new ARMv5e optimizations on classic (non-M) profile
* fix inconsistent ARM_ARCH >= 5 vs == 5
Fixes red in 0c4345475a and ae223933bf
Change-Id: Ieb688679d2d698a19870b1a63b4151a8593e04e0
radix-3, 4 and 5. The gain is bookkeeping: twiddles addressed by
displacement from one base register, post-indexed stores, and C_MUL's
Q15 doubling folded into the add that consumes it.
Modelled: -4.72% ARMv5E, opus_fft_impl -25.3%.
Measured: Clip+ 30.89 -> 29.33 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I173667ebb13308f83bd6fcb4e876f5a56c698cd3
Pre-rotation, post-rotation and mirror, with a packed-twiddle complex
multiply throughout.
Modelled: -8.0% ARMv5E against the C loops.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I3fcaf2343e04e3f3f4462fc442f8fac623729631
deemphasis_stereo_simple on both cores, with the filter state kept
unshifted and shifted inside the add that consumes it.
Modelled: -0.53% ARMv4, -0.96% ARMv5E.
Measured with the three preceding commits: e200v1 49.42 -> 48.96 MHz,
Clip+ 32.83 -> 30.89 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id09297974f3b66a5a193b524d68fd8844f71cc02
ARMv4 has no CLZ, so all 8,680 ilog2 calls went through libgcc's
__clzsi2. Fifteen branchless instructions replace it.
Modelled: -0.91% ARMv4; ARMv5E unaffected, it already emits CLZ.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ib84b97acb687da4101e02bad1b0e8bbd256d32df
comb_filter_const on both cores, and celt_synthesis's SIG_SAT clamp
four samples at a time through one ldm and one stm.
Modelled for the saturation loop: -1.05% ARMv4, -1.42% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I381afab451eb29a91513ce6ab29c1b2d565b7ee9
denormalise_bands on both cores; exp_rotation1, haar1 and the
normalise_residual scaling loop on ARMv5E. Bit-exact.
Modelled: -0.50% ARMv4, -2.91% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: If40f5dbe7cb9e7c12aa0a5b4ac9e73433a850b6e
98.6% of __udivsi3 calls are ec_decode and ec_decode_bin computing
val/ext, where the quotient only matters below 2^16. Exact over 80
million cases including corrupt-stream values.
Modelled: -1.34% ARMv4, -2.37% ARMv5E.
Measured with the previous commit: e200v1 50.80 -> 49.42 MHz,
Clip+ 33.40 -> 32.83 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ia09cb6cd8e42d4885e4d5be0f84862b9b951136c
Cuts realtime decode on the Sansa e200v1 from 52.1 MHz to 50.8 MHz.
clt_mdct_backward was the largest remaining item at 13.5% of decode.
Only the three inner loops move to assembly. The setup stays in C, so
mdct.c remains readable and the assembly needs no knowledge of
mdct_lookup.
What the compiled loops lose is registers. Each needs more live values
than gcc can hold, so it spills the loop-invariant pointers, strides and
limits and reloads them every pass: five stack accesses per iteration in
the post-rotation alone. Holding the twiddle as a 16-bit value and
accumulating the product pair with smull/smlal is what makes the
bookkeeping fit, needing seven live registers where the shifted
MULT16_32_Q15 form needs nine.
ldm/stm helps only where the addressing allows. The post-rotation walks
the buffer from both ends and so reads and writes contiguous pairs. The
pre-rotation reads the spectrum through a runtime stride and writes
through the bitrev table, so only its 8-byte output pair merges, and the
TDAC mirror merges nothing.
Over 160 ms of stereo music, traced under qemu:
clt_mdct_backward 1,037,962 -> 900,982 -13.2%
whole decode 7,695,876 -> 7,558,896 -1.8%
loads 650,157 -> 611,667 -5.9%
stores 350,605 -> 323,605 -7.7%
multiplies 337,493 -> 337,493 unchanged
Accuracy improves substantially, because all three loops keep 32 bits of
each Q15 product where MULT16_32_Q15_armv4 drops the low bit, and the
backward MDCT applies three such rounds per sample. The rounding SNR of
the backward transform rises about 9.5 dB, and its worst case error falls
from 708 to 186. Decoded output differs from the previous build in 90 of
15,360 samples, each by one LSB.
Build with OPUS_ARM_NO_MDCT_ASM to select the C loops instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I3c4404b4dbe581d8bcf1f357266a658a068fdb50
Cuts realtime Opus decode on the Sansa e200v1 (PP502x, ARM7TDMI) from
55.55 MHz to 52.1 MHz.
Two exact changes to celt/kiss_fft.c first:
- kf_bfly5 folds the four cosine products into a shift and a single
multiply. cos(2*pi/5) + cos(4*pi/5) is exactly -1/2, and the Q15
constants satisfy that identity exactly (10126 - 26510 == -16384), so
the substitution gives up no accuracy.
- kf_bfly3, kf_bfly4 and kf_bfly5 peel the pass that twiddles by
twiddles[0], which is 1, by rotating the loop rather than duplicating
the body. In fixed point twiddles[0] is 32767 rather than 32768, so
skipping it also drops a small systematic gain error.
Then celt/arm/kiss_fft_armv4_asm.S replaces the radix-3, radix-4 and
radix-5 bodies, reached through OVERRIDE_kf_bfly3/4/5. The compiled
kernels spill their loop-invariant pointers and reload them every pass,
and gcc will not form ldm/stm from contiguous C accesses: it reorders the
loads while scheduling and does not hand out ascending register pairs.
The assembly keeps the bookkeeping resident and sends the transient
butterfly values to the frame instead, in bursts.
Over 160 ms of stereo music, traced under qemu and costed with an
ARM7TDMI model:
FFT cycles 1,876,926 -> 1,430,748 -23.8%
whole decode 8,142,054 -> 7,695,876 -5.5%
loads 769,935 -> 650,157 -15.6%
stores 409,999 -> 350,605 -14.5%
multiplies 350,909 -> 337,493 -3.8%
text 3,368 -> 2,656 bytes
The radix-5 fold accounts for the whole multiply reduction. The assembly
leaves the count untouched and wins purely on memory traffic.
Accuracy improves by up to 2.9 dB rather than degrading, because the
assembly keeps all 32 bits of each Q15 product where MULT16_32_Q15_armv4
drops the low bit. Decoded output differs from the C build in 27 of
15,360 samples, each by one LSB.
Build with OPUS_ARM_NO_FFT_ASM to select the C butterflies instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I2e8a0904da8a85b654ff648a01234348977a0edc
Several symbols were missing these annotations so the compiler wasn't
handling thumb interworking correctly.
Should resolve tone control crashes seen on clipv1 and our handful of
other 2MB armv5 targets that we build (mostly) in thumb mode.
Change-Id: If8c533a5f10b6592a2d2bec519f2c4d6539d7a66
Most of the churn here occurs because 'filesize' is one of
the redefined filesystem functions, but some of the structs
used by the native FS code also include a 'filesize' member
variable which gets renamed by the macro in some but not all
source files.
It's easier to rename 'filesize()' to 'ffilesize()' rather
than try to clean up the macro mess or renaming the struct
members.
There is weirdness with root_realpath() which now breaks on
native builds because it was assumed to be unprefixed there.
dir_get_info() was unprefixed everywhere but this just seems
inconsistent; make it follow the FS_PREFIX() convention too.
Change-Id: Ic3700c6234ea45f32679c1a8429d70fdb8f4088a
Size the reduced stco lookup with ceiling division. The table stores
chunk entries 0, divider, 2 * divider, and so on, so floor division
allocated one entry too few whenever the original count had a
remainder.
Treat a valid media-information atom whose first child is not smhd as
a non-audio track and skip the remainder. Continue scanning subsequent
tracks so AAC decoding works when an MP4 places its video track before
the audio track, while still rejecting malformed atom sizes and
malformed sound headers.
Keep these container fixes independent of the H.264 player and target
driver so they can be reviewed and applied to libm4a on their own.
Build-tested as part of the normal and isolated-runtime iPod 6G
configurations and hardware-tested with AAC audio in M4V playback.
Change-Id: I8c711525932f54ecbdad982c7f7ddc9490c5668d
This commit does the following changes to the 3ds port:
- Rename the target from ctru to 3ds.
- Rename all files and functions with the ctru naming convention to 3ds.
- Created a new file and folder structure that will better integrate future console ports that share the same codebase.
- Fixed a buffer overflow bug in pcm code.
Change-Id: I17c6f86df64eb99dd2b653485d70832ff46b2ba8
Allows any EQ band to be set to any of Low Shelf, Peak, or High Shelf, instead of hardcoding the types per band.
Change-Id: I470ab916359092ba465e7b6331baed3bf11b2fc9
Use flac_seek by time even when elapsedtime is 0, and apply it as a fallback for failed offset seeks since it provides more robust error recovery.
Change-Id: I438888fba02bda38137f3f1347bb1f657ef9c166
skin_buffer_to_offset() call can never return a "negative" pointer
(since it just returns the pointer as-is) so don't bother to check.
Change-Id: Id86d53abd7ab1fb071ca54421ebe3b5ff2981c02
add get_metadata_afmt function so we don't have to extra functions
remove unneeded bounds check on audio_format in rbcodec_format_is_atomic()
add bounds check on audio_format in get_metadata_afmt()
Change-Id: I76bd869100b000579c6546f0670ba4ba2c541f22
tagcache.c add_tagcache() and potentially
skin_tokens.c wps_playlist_percent_prepare()
make calls to probe_file_format() prior to calling get_metadata_ex
resulting in some small amout of duplicated work
especially in the case of add_tagcache this can add
up to a lot of duplicated work
breaks out audio_fmt so these can supply the afmt other callers just
supply probe_file_format(trackname) in the function call
Change-Id: I8084213b8ee7e04d76dce0986beb83d443ac804b