Converting 48 kHz to 44.1 kHz: Aliasing, High-Frequency Loss and Peaks of swr and soxr, Measured

Contents
- How the measurement was made
- Passband: ffmpeg’s default already attenuates below 20 kHz
- Above 22.05 kHz: aliasing between −9.5 and below −180 dB
- Peaks: the true peak stayed, the sample peak rose by up to 1.3 dB
- Pre-ringing and processing time: hardly any differences
- What this means for the export
- Context and limits
- Questions and answers
- Sources
FL Studio exports at the project’s sample rate. The export window has no setting of its own for it; according to the manual, it belongs in the Audio Settings. When a project runs at 48 kHz and a 44.1 kHz file is needed, a resampler recalculates the samples. When the sampling rate is lowered, it first has to remove whatever lies above half the new sampling rate, at 44.1 kHz everything above 22.05 kHz: Julius O. Smith describes how the lowpass cutoff moves below half the new, lower sampling rate for this purpose. Whatever remains above it folds back into the range below as aliasing.
The measurement compares the two resamplers that ffmpeg ships with: its own swr with three filter lengths, including the default, and the SoX Resampler soxr at two quality levels. Sine tones, a single impulse and two loud masters from the test mix of the measurement True Peak After Encoding were measured, this time generated at 48 kHz.

How the measurement was made
- Methods: ffmpeg 7.1.5, aresample filter with 44.1 kHz output in 32-bit float. swr with filter_size 16, 32 (default) and 128, soxr with precision 20 (default, SoX “High Quality” according to the documentation) and 28 (“Very High Quality”). The cutoff stayed at the default everywhere: for swr 0.97 of half the sampling rate as the 6 dB point, for soxr 0.91 as the 0 dB point.
- Checks:
-ar 44100without further options produced the same sample values as aresample with default settings. Without the NEON routines of the ARM computer (-cpuflags 0), the largest deviation for swr was 131 dB below the level of the tones; soxr showed none. - Sine tones: one second each at −6 dBFS at 48 kHz, analysed over the middle half second with a least-squares fit. Tones from 1 to 22 kHz show the passband level. Tones from 22.5 to 23.9 kHz lie above the new Nyquist frequency of 22.05 kHz; the measurement captured what arrives at 44.1 kHz minus the tone frequency.
- Single impulse: duration of the pre-ringing before the main peak down to −60 and −100 dB.
- Masters: the test mix, regenerated at 48 kHz, once with a limiter at −9 LUFS and −1 dBTP, once hard-clipped at −1 dBFS to −8 LUFS. The hi-hats made from filtered noise reach up to 24 kHz; the share above 20 kHz was 28.5 and 27.3 dB below the total energy. True peak per ITU-R BS.1770-5 and a 16-times cross-check, plus sample peak and samples above 0 dBFS.
- Processing time: the limiter master converted three times per method on a Raspberry Pi 4, counting the shortest time including reading and writing the files.
Passband: ffmpeg’s default already attenuates below 20 kHz
| Method | 18 kHz | 19 kHz | 20 kHz | 20.5 kHz | 21 kHz |
|---|---|---|---|---|---|
| swr, filter_size 16 | −0.81 dB | −1.67 dB | −3.03 dB | −3.95 dB | −5.04 dB |
| swr, default (filter_size 32) | −0.01 dB | −0.22 dB | −1.28 dB | −2.43 dB | −4.18 dB |
| swr, filter_size 128 | 0.00 dB | 0.00 dB | 0.00 dB | 0.00 dB | −1.01 dB |
| soxr, default (precision 20) | 0.00 dB | 0.00 dB | −0.01 dB | −0.07 dB | −3.75 dB |
| soxr, precision 28 | 0.00 dB | 0.00 dB | 0.00 dB | −0.03 dB | −4.06 dB |
With default settings, swr lost 0.22 dB at 19 kHz and 1.28 dB at 20 kHz; with filter_size 16 it was already 3.03 dB at 20 kHz. With filter_size 128 and with both soxr levels, the level stayed within 0.01 dB up to 20 kHz and within 0.07 dB up to 20.5 kHz; the steep drop only begins above that.
Above 22.05 kHz: aliasing between −9.5 and below −180 dB
| Method | Tone 22.5 kHz, alias at 21.6 kHz | Tone 23 kHz, alias at 21.1 kHz | Tone 23.5 kHz, alias at 20.6 kHz | Tone 23.9 kHz, alias at 20.2 kHz |
|---|---|---|---|---|
| swr, filter_size 16 | −9.5 dB | −11.5 dB | −13.8 dB | −15.8 dB |
| swr, default (filter_size 32) | −14.3 dB | −19.9 dB | −27.1 dB | −34.4 dB |
| swr, filter_size 128 | −96.2 dB | −94.7 dB | −104.0 dB | −97.9 dB |
| soxr, default (precision 20) | −136.6 dB | −148.6 dB | −141.6 dB | −162.2 dB |
| soxr, precision 28 | below −180 dB | below −180 dB | below −180 dB | below −180 dB |
A tone at 22.5 kHz cannot be represented at 44.1 kHz; if it remains, it appears mirrored at 21.6 kHz. The swr default attenuated it by only 14.3 dB, filter_size 16 by 9.5 dB. Only further away from the Nyquist frequency did the attenuation grow, for the default up to 34.4 dB at 23.9 kHz. swr with filter_size 128 reached 94.7 to 104.0 dB, soxr 136.6 to 162.2 dB and with precision 28 more than 180 dB. Smith names the price of steep filters: for a given stop-band specification, a lowpass filter needs approximately twice as many calculations per sample for each halving of the transition band width.
Peaks: the true peak stayed, the sample peak rose by up to 1.3 dB
| Master | True peak 4 times | True peak 16 times | Sample peak | Samples above 0 dBFS |
|---|---|---|---|---|
| Limiter, −9 LUFS, −1 dBTP: Master at 48 kHz | −1.05 dBTP | −0.79 dBTP | −2.31 dBFS | 0 |
| Limiter: swr, filter_size 16 | −1.16 dBTP | −1.07 dBTP | −1.53 dBFS | 0 |
| Limiter: swr, default (filter_size 32) | −1.03 dBTP | −0.96 dBTP | −1.44 dBFS | 0 |
| Limiter: swr, filter_size 128 | −1.04 dBTP | −0.98 dBTP | −1.35 dBFS | 0 |
| Limiter: soxr, default (precision 20) | −1.05 dBTP | −0.95 dBTP | −1.40 dBFS | 0 |
| Limiter: soxr, precision 28 | −1.05 dBTP | −0.95 dBTP | −1.41 dBFS | 0 |
| Hard clipping at −1 dBFS, −8 LUFS: Master at 48 kHz | +0.38 dBTP | +0.72 dBTP | −1.00 dBFS | 0 |
| Hard clipping at −1 dBFS: swr, filter_size 16 | +0.28 dBTP | +0.38 dBTP | +0.01 dBFS | 2 |
| Hard clipping at −1 dBFS: swr, default (filter_size 32) | +0.41 dBTP | +0.54 dBTP | +0.18 dBFS | 18 |
| Hard clipping at −1 dBFS: swr, filter_size 128 | +0.47 dBTP | +0.64 dBTP | +0.27 dBFS | 41 |
| Hard clipping at −1 dBFS: soxr, default (precision 20) | +0.45 dBTP | +0.64 dBTP | +0.26 dBFS | 42 |
| Hard clipping at −1 dBFS: soxr, precision 28 | +0.44 dBTP | +0.64 dBTP | +0.26 dBFS | 44 |
Loudness stayed the same for every conversion (resolution 0.1 LU). The true peak measured 4 times changed by no more than 0.11 dB; in the 16-times cross-check it fell by 0.08 to 0.34 dB, most with filter_size 16, which attenuates the highs most. The samples, by contrast, hit different points of the waveform after conversion. For the limiter master, the sample peak rose from −2.31 to as much as −1.35 dBFS. The hard-clipped master had no sample above −1.00 dBFS at 48 kHz; after conversion, 2 to 44 samples were above 0 dBFS, 42 and 44 with soxr. In a 16-bit or 24-bit file, such values are clipped. The limit on the samples at 48 kHz therefore does not protect the 44.1 kHz file, whereas the true peak of the 48 kHz file already showed beforehand how high the values can reach.
Pre-ringing and processing time: hardly any differences
| Method | Processing time for 32 s of stereo | faster than real time | Pre-ringing down to −60 dB | down to −100 dB |
|---|---|---|---|---|
| swr, filter_size 16 | 0.50 s | 64 times | 0.14 ms | 0.18 ms |
| swr, default (filter_size 32) | 0.55 s | 58 times | 0.27 ms | 0.36 ms |
| swr, filter_size 128 | 0.81 s | 39 times | 0.68 ms | 1.41 ms |
| soxr, default (precision 20) | 0.70 s | 46 times | 0.98 ms | 1.93 ms |
| soxr, precision 28 | 0.79 s | 40 times | 1.32 ms | 2.27 ms |
All five methods rang for the same length before and after the impulse, as linear-phase filters do. Down to −60 dB, the pre-ringing lasted from 0.14 ms with filter_size 16 to 1.32 ms with soxr at precision 28, down to −100 dB no more than 2.27 ms. The more accurate methods cost little time on the Raspberry Pi 4: 32 seconds of stereo took 0.55 s with the swr default, 0.70 s with soxr and 0.79 s with precision 28, each including reading and writing the files.
What this means for the export
- For a 44.1 kHz file from a 48 kHz export,
-af aresample=44100:resampler=soxrchanged the level by no more than 0.01 dB up to 20 kHz and attenuated content above 22.05 kHz by at least 137 dB;:precision=28cost 0.1 s more. -ar 44100on its own uses the swr default: 1.3 dB less at 20 kHz and only 14 dB of attenuation for content just above 22.05 kHz. Within swr, filter_size=128 brought the deviation up to 20 kHz below 0.01 dB and the attenuation to 95 to 104 dB.- What matters for the peaks after conversion is the true peak of the 48 kHz file, not its sample peak. The LUFS and True Peak Meter measures both in the browser.
- Conversion to 16-bit with dither follows as the last step after the sample rate conversion; see Dither When Exporting to 16-Bit.
- The FL Studio manual describes the same principle for interpolation when transposing samples: linear interpolation can cause aliasing, and for the final render it recommends settings above 64-point sinc if aliasing is audible with lower methods. The setting is covered in the mastering tutorial for FL Studio.
Context and limits
The measurement covered the resamplers in ffmpeg, not the conversion in FL Studio or other programs. The test mix is synthetic; its hi-hats made from filtered noise reach up to 24 kHz, while many recordings contain less above 20 kHz. A listening test was not part of the measurement: the folded content lands between 20.2 and 21.6 kHz, and whether it becomes audible there depends on material, hearing and playback. The processing times apply to the Raspberry Pi 4 used for the measurement; the absolute values change on other hardware.
Questions and answers
Why does the more accurate resampler produce more samples above 0 dBFS on the hard-clipped master?
Because it reproduces the waveform between the samples more faithfully. Hard clipping at −1 dBFS creates edges with content reaching up to the top of the band, and the waveform described by these samples swings beyond the clipping level between them: the true peak of the 48 kHz file was +0.72 dBTP measured with 16-times oversampling, although no sample exceeded −1.00 dBFS.
A resampler calculates new samples at different points in time, so it reads this waveform at new positions and also hits spots close to the peaks. The flatter its passband, the more of the high frequencies that make up the overshoots survive. filter_size 16 already attenuates by 3.03 dB at 20 kHz, smoothing the peaks, and ended up with 2 samples above 0 dBFS; soxr and filter_size 128 leave the high end intact and ended up with 41 to 44.
The problem therefore lies in the master, not in the resampler. What matters is the true peak of the 48 kHz file: if the file is lowered until that true peak sits below 0 dBTP with a little margin, or if limiting only happens after the conversion, the samples of the 44.1 kHz file stay below 0 dBFS, at 16 or 24 bits as well.
Can the aliasing of swr also be fixed with a lower cutoff instead of a larger filter_size?
Only at the expense of the high end. The cutoff, for swr a 6 dB point at 0.97 of half the sampling rate, around 21.4 kHz, determines where the transition band lies; filter_size determines how wide it is. With the default of 32 it is so wide that it reaches both below 20 kHz and beyond 22.05 kHz: hence the 1.28 dB loss at 20 kHz and only 14.3 dB of attenuation at 22.5 kHz.
A lower cutoff merely shifts this wide band downward. The attenuation just above 22.05 kHz then rises, but the loss below 20 kHz grows along with it; this was not measured, it follows from where the transition band lies. The transition only becomes narrower with a longer filter, and that is exactly Smith’s rule of thumb: every halving of the transition band costs roughly twice as many calculations per sample. On the Raspberry Pi used for the measurement that price was small, with filter_size 128 taking 0.81 instead of 0.55 s for 32 seconds of stereo.
Does it matter whether the conversion to 16-bit comes before or after the conversion to 44.1 kHz?
Yes. The resampler calculates new samples that no longer sit on the 16-bit grid, so its output has to be rounded a second time, either without suitable dither or with a second layer of noise. On top of that come the samples above 0 dBFS that the conversion can produce and that are clipped in a 16-bit file. That is why the conversion to 16-bit with dither comes last.
Sources
- FFmpeg Resampler Documentation
- ffmpeg Documentation
- Implementation – Digital Audio Resampling Home Page (Julius O. Smith III)
- Theory of Ideal Bandlimited Interpolation – Digital Audio Resampling Home Page (Julius O. Smith III)
- Exporting Audio & MIDI – FL Studio Online Manual
- ITU-R BS.1770: Algorithms to measure audio programme loudness and true-peak audio level