The 8, 12 and 16 kHz to 48 kHz steps visit only two or three FIR phases
in a fixed cycle, so each set of input samples is read once and reused
across outputs. Bit-exact; OPUS_NO_SILK_FIXED_PHASE disables it.
Measured with silk_5.opus (WB SILK):
e200v1: 36.31 -> 33.38 MHz, -8.1%
Clip+: 21.66 -> 20.93 MHz, -3.4%
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: Icb0f8ded62739dfee1f574e4888a54c778aa3d53
A total sample count of 0 means unknown. It made the track length 0 and the
bitrate estimate in flac_init() divided by it, crashing the codec (FLAC
decoder testbench file 45). Report a bitrate of 0 in that case.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
An escape code with a raw bit width of 0 means every residual in the
partition is zero. get_sbits(&gb, 0) shifts by 32, which is undefined and
returned stale cache bits instead of 0, so such streams decoded to garbage
(FLAC decoder testbench file 64).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
print_mp3entry() dereferences mb_track_id, but get_metadata() does not set
every field of the uninitialized stack struct. For FLAC files this left
garbage in the pointer and warble segfaulted in strlen about a third of the
time, before decoding started. Clear the struct first.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Upstream's OPUS_ARM_INLINE_ASM means any ARM with inline assembly, with
OPUS_ARM_INLINE_EDSP layered on top for ARMv5E. config.h instead defines
exactly one of them per core, so the names read as broader than they are,
and code added ARM_ARCH tests beside them to pin the scope down. Renamed
to OPUS_ARM_ASM_ARMV4_ONLY and OPUS_ARM_ASM_ARMV5E_AND_LATER throughout
celt and silk, upstream files included; README.rockbox records it for the
next libopus sync.
No code change: opus.elf disassembly and section sizes are identical
before and after on ARMv4 (e200v1), ARMv5E (Clip+) and ARMv6 (iPod Nano
4G).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Change-Id: I89932d348dedd76748a1bc7b3f2e7de1c5be49c8
Four `ARM_ARCH == 5` checks -- in SOURCES and celt/arm/{bands_arm,
comb_filter_arm,vq_arm}.h -- excluded ARMv6 from every ARMv5E kernel:
denorm_band, haar1, comb_filter_const, celt_sat, deemphasis_stereo_simple,
exp_rotation1 and normres_scale all silently fell back to plain C on
ARM1136/ARM1176, since ARMv6 is a strict superset of the EDSP instructions
those kernels use. Widened to ARM_ARCH >= 5, matching config.h's own
OPUS_ARM_INLINE_EDSP gate, which was already ARM_ARCH > 4.
These four can't be fixed at the commits that introduced them: those
commits are already merged into master under different SHAs. A fifth
instance of the same bug, in celt/arm/mdct_armv5e.h, was fixed at its
origin commit instead, since that one is still open for review.
Verified on both native ARMv6 targets (iPod Nano 4G, ARM1176JZ-S; Gigabeat
S, ARM1136JF-S) and the hosted Samsung YP-R0 (ARM1176JZ-S, cross toolchain
built for the occasion): all three now link and call all 16 ARMv5E
kernels, where they linked and called zero before this fix. Decoded PCM
bit-identical to the ARMv5E build under qemu.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Change-Id: I7e6f6b421b9c0203812fb45d2258c7eb81738d98
53% of the dimensions a decode walks hold no pulses, and that arm leaves
_k alone, so the two CELT_PVQ_U_ROW pointers stay valid. U is
non-decreasing in _k, so p <= _i < q is the single unsigned test
(_i-p) < (q-p). Bit-exact over 8.2M samples.
Modelled: -0.40% ARMv4, -0.78% ARMv5E; the function -4.8% and -7.1%.
Measured: e200v1 39.19 -> 38.96 MHz, Clip+ 28.13 -> 28.08 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id50be222e101aeabd9a1f4cf5380d3244da610ff
At stride 1 the rotation chain writes X[i+stride] and reads it straight
back, so one load an iteration is redundant and one store is dead. The
kernel carries that value, narrowed, since mul reads all 32 bits where
smulbb does not. Bit-exact over 8.2M samples.
Modelled: -0.87% ARMv4, the function -10.0%.
Measured: e200v1 39.59 -> 39.19 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Idbd1a088bc805fecfb9ee54fc372ad16a7d91606
34 functions and the ARMv4 kernels, chosen by a greedy fill of the free
codec IRAM window ranked by cycles per byte. The mixed-radix
butterflies give their IRAM back, being unreachable under the prime
factor transform.
The cycle model does not see this at all: it models no instruction cache.
Measured: e200v1 48.96 -> 42.60 MHz; reclaiming the mixed-radix IRAM was
a further -0.52%.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I9e22025cba99521750a3b664cd0d59357fcd0810
Every 48 kHz CELT length is 15 times a power of two and the factors are
co-prime, so the inter-stage twiddles -- 73% of the FFT multiplies at
N=480 -- vanish. New pfa_fft15 and mdct_postrot_pfa kernels on both
cores; accuracy also improves 0.3 to 0.4 dB against opusdec.
Modelled: -6.17% ARMv4, -1.82% ARMv5E.
Measured: e200v1 42.33 -> 40.10 MHz, Clip+ 29.30 -> 28.24 MHz.
Rescheduling these kernels, folded in here, measured a further -0.75% on
e200v1 and -0.39% on Clip+.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ifa4a30d1045e905522828958949954615faa440d
One ldr fetches two celt_norm coefficients and smlabb/smlatt take the
halves apart; scalar fallback when the three pointers disagree on
alignment.
Modelled: -0.88% ARMv5E.
Measured: Clip+ 28.13 -> 28.08 MHz, the same build with and without it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I79719132ccecfb7664616bca019e3856ee3479da
Rockbox lists no encoder, so CELT_DECODE_ONLY folds away the thirteen
encoder branches in celt/bands.c and stops gcc keeping their values live
across quant_partition's recursive calls. 2,976 bytes smaller.
Modelled: -0.47% ARMv4, -0.69% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ie96a0a63e035ab0cdcdbe9eeeb6691f24fecfa41
Its reachable input set is 2,985 (qn,i) pairs for any stream ever, so it
tabulates exactly in 6,484 bytes. Verified exhaustively against the
compiled functions. Only built for PP5022/PP5024, where the tables fit
the 80 KB IRAM window; elsewhere it measured no gain.
Measured, e200v1: 39.18 -> 38.96 MHz with the tables in IRAM, 39.07 with
them in DRAM. Clip+: 28.05 -> 28.08 MHz, so not built there.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I38417793dc337ec6cb61139cc919d23ac09b04dd
This re-enables upstream ASM optimizations for Cortex-M. Only the
our downstream improvements are disabled, as they do not assemble in
thumb2 mode.
Change-Id: Icca1b3dbf04786c7714fb4eef92ad66aa55132f3
* Only use new ARMv5e optimizations on classic (non-M) profile
* fix inconsistent ARM_ARCH >= 5 vs == 5
Fixes red in 0c4345475a and ae223933bf
Change-Id: Ieb688679d2d698a19870b1a63b4151a8593e04e0
radix-3, 4 and 5. The gain is bookkeeping: twiddles addressed by
displacement from one base register, post-indexed stores, and C_MUL's
Q15 doubling folded into the add that consumes it.
Modelled: -4.72% ARMv5E, opus_fft_impl -25.3%.
Measured: Clip+ 30.89 -> 29.33 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I173667ebb13308f83bd6fcb4e876f5a56c698cd3
Pre-rotation, post-rotation and mirror, with a packed-twiddle complex
multiply throughout.
Modelled: -8.0% ARMv5E against the C loops.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I3fcaf2343e04e3f3f4462fc442f8fac623729631
deemphasis_stereo_simple on both cores, with the filter state kept
unshifted and shifted inside the add that consumes it.
Modelled: -0.53% ARMv4, -0.96% ARMv5E.
Measured with the three preceding commits: e200v1 49.42 -> 48.96 MHz,
Clip+ 32.83 -> 30.89 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Id09297974f3b66a5a193b524d68fd8844f71cc02
ARMv4 has no CLZ, so all 8,680 ilog2 calls went through libgcc's
__clzsi2. Fifteen branchless instructions replace it.
Modelled: -0.91% ARMv4; ARMv5E unaffected, it already emits CLZ.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ib84b97acb687da4101e02bad1b0e8bbd256d32df
comb_filter_const on both cores, and celt_synthesis's SIG_SAT clamp
four samples at a time through one ldm and one stm.
Modelled for the saturation loop: -1.05% ARMv4, -1.42% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I381afab451eb29a91513ce6ab29c1b2d565b7ee9
denormalise_bands on both cores; exp_rotation1, haar1 and the
normalise_residual scaling loop on ARMv5E. Bit-exact.
Modelled: -0.50% ARMv4, -2.91% ARMv5E.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: If40f5dbe7cb9e7c12aa0a5b4ac9e73433a850b6e
98.6% of __udivsi3 calls are ec_decode and ec_decode_bin computing
val/ext, where the quotient only matters below 2^16. Exact over 80
million cases including corrupt-stream values.
Modelled: -1.34% ARMv4, -2.37% ARMv5E.
Measured with the previous commit: e200v1 50.80 -> 49.42 MHz,
Clip+ 33.40 -> 32.83 MHz.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: Ia09cb6cd8e42d4885e4d5be0f84862b9b951136c
Cuts realtime decode on the Sansa e200v1 from 52.1 MHz to 50.8 MHz.
clt_mdct_backward was the largest remaining item at 13.5% of decode.
Only the three inner loops move to assembly. The setup stays in C, so
mdct.c remains readable and the assembly needs no knowledge of
mdct_lookup.
What the compiled loops lose is registers. Each needs more live values
than gcc can hold, so it spills the loop-invariant pointers, strides and
limits and reloads them every pass: five stack accesses per iteration in
the post-rotation alone. Holding the twiddle as a 16-bit value and
accumulating the product pair with smull/smlal is what makes the
bookkeeping fit, needing seven live registers where the shifted
MULT16_32_Q15 form needs nine.
ldm/stm helps only where the addressing allows. The post-rotation walks
the buffer from both ends and so reads and writes contiguous pairs. The
pre-rotation reads the spectrum through a runtime stride and writes
through the bitrev table, so only its 8-byte output pair merges, and the
TDAC mirror merges nothing.
Over 160 ms of stereo music, traced under qemu:
clt_mdct_backward 1,037,962 -> 900,982 -13.2%
whole decode 7,695,876 -> 7,558,896 -1.8%
loads 650,157 -> 611,667 -5.9%
stores 350,605 -> 323,605 -7.7%
multiplies 337,493 -> 337,493 unchanged
Accuracy improves substantially, because all three loops keep 32 bits of
each Q15 product where MULT16_32_Q15_armv4 drops the low bit, and the
backward MDCT applies three such rounds per sample. The rounding SNR of
the backward transform rises about 9.5 dB, and its worst case error falls
from 708 to 186. Decoded output differs from the previous build in 90 of
15,360 samples, each by one LSB.
Build with OPUS_ARM_NO_MDCT_ASM to select the C loops instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I3c4404b4dbe581d8bcf1f357266a658a068fdb50
Cuts realtime Opus decode on the Sansa e200v1 (PP502x, ARM7TDMI) from
55.55 MHz to 52.1 MHz.
Two exact changes to celt/kiss_fft.c first:
- kf_bfly5 folds the four cosine products into a shift and a single
multiply. cos(2*pi/5) + cos(4*pi/5) is exactly -1/2, and the Q15
constants satisfy that identity exactly (10126 - 26510 == -16384), so
the substitution gives up no accuracy.
- kf_bfly3, kf_bfly4 and kf_bfly5 peel the pass that twiddles by
twiddles[0], which is 1, by rotating the loop rather than duplicating
the body. In fixed point twiddles[0] is 32767 rather than 32768, so
skipping it also drops a small systematic gain error.
Then celt/arm/kiss_fft_armv4_asm.S replaces the radix-3, radix-4 and
radix-5 bodies, reached through OVERRIDE_kf_bfly3/4/5. The compiled
kernels spill their loop-invariant pointers and reload them every pass,
and gcc will not form ldm/stm from contiguous C accesses: it reorders the
loads while scheduling and does not hand out ascending register pairs.
The assembly keeps the bookkeeping resident and sends the transient
butterfly values to the frame instead, in bursts.
Over 160 ms of stereo music, traced under qemu and costed with an
ARM7TDMI model:
FFT cycles 1,876,926 -> 1,430,748 -23.8%
whole decode 8,142,054 -> 7,695,876 -5.5%
loads 769,935 -> 650,157 -15.6%
stores 409,999 -> 350,605 -14.5%
multiplies 350,909 -> 337,493 -3.8%
text 3,368 -> 2,656 bytes
The radix-5 fold accounts for the whole multiply reduction. The assembly
leaves the count untouched and wins purely on memory traffic.
Accuracy improves by up to 2.9 dB rather than degrading, because the
assembly keeps all 32 bits of each Q15 product where MULT16_32_Q15_armv4
drops the low bit. Decoded output differs from the C build in 27 of
15,360 samples, each by one LSB.
Build with OPUS_ARM_NO_FFT_ASM to select the C butterflies instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I2e8a0904da8a85b654ff648a01234348977a0edc
Several symbols were missing these annotations so the compiler wasn't
handling thumb interworking correctly.
Should resolve tone control crashes seen on clipv1 and our handful of
other 2MB armv5 targets that we build (mostly) in thumb mode.
Change-Id: If8c533a5f10b6592a2d2bec519f2c4d6539d7a66
Most of the churn here occurs because 'filesize' is one of
the redefined filesystem functions, but some of the structs
used by the native FS code also include a 'filesize' member
variable which gets renamed by the macro in some but not all
source files.
It's easier to rename 'filesize()' to 'ffilesize()' rather
than try to clean up the macro mess or renaming the struct
members.
There is weirdness with root_realpath() which now breaks on
native builds because it was assumed to be unprefixed there.
dir_get_info() was unprefixed everywhere but this just seems
inconsistent; make it follow the FS_PREFIX() convention too.
Change-Id: Ic3700c6234ea45f32679c1a8429d70fdb8f4088a
Size the reduced stco lookup with ceiling division. The table stores
chunk entries 0, divider, 2 * divider, and so on, so floor division
allocated one entry too few whenever the original count had a
remainder.
Treat a valid media-information atom whose first child is not smhd as
a non-audio track and skip the remainder. Continue scanning subsequent
tracks so AAC decoding works when an MP4 places its video track before
the audio track, while still rejecting malformed atom sizes and
malformed sound headers.
Keep these container fixes independent of the H.264 player and target
driver so they can be reviewed and applied to libm4a on their own.
Build-tested as part of the normal and isolated-runtime iPod 6G
configurations and hardware-tested with AAC audio in M4V playback.
Change-Id: I8c711525932f54ecbdad982c7f7ddc9490c5668d
This commit does the following changes to the 3ds port:
- Rename the target from ctru to 3ds.
- Rename all files and functions with the ctru naming convention to 3ds.
- Created a new file and folder structure that will better integrate future console ports that share the same codebase.
- Fixed a buffer overflow bug in pcm code.
Change-Id: I17c6f86df64eb99dd2b653485d70832ff46b2ba8
Allows any EQ band to be set to any of Low Shelf, Peak, or High Shelf, instead of hardcoding the types per band.
Change-Id: I470ab916359092ba465e7b6331baed3bf11b2fc9
Use flac_seek by time even when elapsedtime is 0, and apply it as a fallback for failed offset seeks since it provides more robust error recovery.
Change-Id: I438888fba02bda38137f3f1347bb1f657ef9c166
skin_buffer_to_offset() call can never return a "negative" pointer
(since it just returns the pointer as-is) so don't bother to check.
Change-Id: Id86d53abd7ab1fb071ca54421ebe3b5ff2981c02
add get_metadata_afmt function so we don't have to extra functions
remove unneeded bounds check on audio_format in rbcodec_format_is_atomic()
add bounds check on audio_format in get_metadata_afmt()
Change-Id: I76bd869100b000579c6546f0670ba4ba2c541f22
tagcache.c add_tagcache() and potentially
skin_tokens.c wps_playlist_percent_prepare()
make calls to probe_file_format() prior to calling get_metadata_ex
resulting in some small amout of duplicated work
especially in the case of add_tagcache this can add
up to a lot of duplicated work
breaks out audio_fmt so these can supply the afmt other callers just
supply probe_file_format(trackname) in the function call
Change-Id: I8084213b8ee7e04d76dce0986beb83d443ac804b
%pP reports playlist progress by position index, which treats every
track as equally long. For playlists with tracks of unequal length --
audiobooks with chapters anywhere from two minutes to an hour are the
motivating case -- position is a poor proxy for listening progress.
%pX reports the played percentage of the whole playlist by time: the
summed length of all preceding tracks plus the elapsed time in the
current one, relative to the playlist's total duration. It can be
used as a value, in a conditional, with %if(), or as a bar tag like
%pb.
If a playlist contains more than 500 tracks or the scan is taking too
long and the user aborts the tag will fallback to the behavior of %pP
except the progress through the current track will be included in the
returned percentage
--------------------------------------
Computing this needs every track's length, and reading metadata for
every track is too slow and disk-heavy for a tag that refreshes on the
WPS. Instead each track's length is estimated from its file size: the
skin engine scans the playlist in the background, a batch of files each
skin refresh, opening each file only to read its size (directory
metadata, no header parse). Size is turned into time by calibrating one
file of each type -- the first file of each extension is parsed once
with get_metadata to learn its bytes-per-second, and every later file
of that type reuses it. A single-format playlist, the usual audiobook
case, parses exactly one file and stat's the rest.
The result is an estimate -- bitrate varies within a type, especially
for VBR -- but it is cheap and accurate enough for a progress
indicator, and the playing track always contributes its exact elapsed
time. Per-track lengths are stored as two bytes of minutes each in a
movable buffer sized to the track count and allocated only while the
tag is in use; the cache is keyed on the track count and a crc of the
first, middle and last filenames, so a playlist swap or reshuffle is
caught. If the buffer cannot be allocated (a very large playlist on a
low-memory target) the tag falls back to position-based progress.
The buffer is allocated when playback starts (and when the now-playing
screen is opened with playback already active) rather than lazily on the
first WPS refresh, so the one-time allocation happens at a playback
boundary instead of during steady-state playback; it falls back to
allocating on first use if that point is missed.
Until the scan finishes the tag does not blank: it returns an instant
equal-weight estimate -- (completed tracks + fraction through the
current one) / track count -- which sharpens into the size-based value
as the scan fills in. So a theme can just use %pX and it is correct
from the first frame, and %?pX is true whenever a playlist is loaded.
Tested on a Sansa Clip Zip and in the simulator. On a 216-track,
~745 MB single-format audiobook playlist the length scan dropped from
~4150 ms (get_metadata on every track) to ~174 ms (one parse plus 215
file-size stats), roughly 24x lighter.
Change-Id: I6e572e78a10444bd513ddc77e30da04aa5153ef2
currently logf() in codec are printed unconditionally when
ROCKBOX_HAS_LOGF, which is confusing.
Gate them behind LOGF_ENABLE to align with the main binary.
Change-Id: I4eb44604000acded55d0af8869145c4f3d77efcb
clang defines both of __LITTLE_ENDIAN__ and __mips__, causing
confusion in the byte order detection.
Change-Id: If95255974e9e3e5d0554410bbab65bc2af1bb1c7
Load stts (time to sample) table on demand if it doesn't fit in RAM.
Fixes FS#13889 (in most cases I've seen files with single element in this table and problematic one has 111200 elements)
Change-Id: I719f92a4512a45472739587e81861b9bc545f349
GCC16 is stricter about reporting these, and while we can rely on
the compiler to optimize away dead stores, we're better off eliminating
them from the source code.
Change-Id: I14570a986811a77ca656c60d792593ff8c458571