LW IT Solutions
« Blog Overview /Music Production / Converting 48 kHz to 44.1 kHz: Aliasing,...

Converting 48 kHz to 44.1 kHz: Aliasing, High-Frequency Loss and Peaks of swr and soxr, Measured

Converting 48 kHz to 44.1 kHz: Aliasing, High-Frequency Loss and Peaks of swr and soxr, Measured
Contents
  1. How the measurement was made
  2. Passband: ffmpeg’s default already attenuates below 20 kHz
  3. Above 22.05 kHz: aliasing between −9.5 and below −180 dB
  4. Peaks: the true peak stayed, the sample peak rose by up to 1.3 dB
  5. Pre-ringing and processing time: hardly any differences
  6. What this means for the export
  7. Context and limits
  8. Questions and answers
  9. Sources

FL Studio exports at the project’s sample rate. The export window has no setting of its own for it; according to the manual, it belongs in the Audio Settings. When a project runs at 48 kHz and a 44.1 kHz file is needed, a resampler recalculates the samples. When the sampling rate is lowered, it first has to remove whatever lies above half the new sampling rate, at 44.1 kHz everything above 22.05 kHz: Julius O. Smith describes how the lowpass cutoff moves below half the new, lower sampling rate for this purpose. Whatever remains above it folds back into the range below as aliasing.

The measurement compares the two resamplers that ffmpeg ships with: its own swr with three filter lengths, including the default, and the SoX Resampler soxr at two quality levels. Sine tones, a single impulse and two loud masters from the test mix of the measurement True Peak After Encoding were measured, this time generated at 48 kHz.

Two line charts for five methods. On the left, the level of sine tones between 15 and 22 kHz after conversion: swr with filter_size 16 already drops from 18 kHz and is 3 dB lower at 20 kHz, the swr default 1.3 dB, swr with filter_size 128 and both soxr levels stay close to 0 dB up to 20.5 kHz and only then fall steeply. On the right, the level of the folded content for tones between 22.5 and 23.9 kHz: swr with filter_size 16 at −10 to −16 dB, the default at −14 to −34 dB, filter_size 128 around −100 dB, soxr between −137 and −162 dB, soxr with precision 28 below −180 dB. Four key figures below
Passband level and aliasing when converting from 48 to 44.1 kHz, per method.

How the measurement was made

  • Methods: ffmpeg 7.1.5, aresample filter with 44.1 kHz output in 32-bit float. swr with filter_size 16, 32 (default) and 128, soxr with precision 20 (default, SoX “High Quality” according to the documentation) and 28 (“Very High Quality”). The cutoff stayed at the default everywhere: for swr 0.97 of half the sampling rate as the 6 dB point, for soxr 0.91 as the 0 dB point.
  • Checks: -ar 44100 without further options produced the same sample values as aresample with default settings. Without the NEON routines of the ARM computer (-cpuflags 0), the largest deviation for swr was 131 dB below the level of the tones; soxr showed none.
  • Sine tones: one second each at −6 dBFS at 48 kHz, analysed over the middle half second with a least-squares fit. Tones from 1 to 22 kHz show the passband level. Tones from 22.5 to 23.9 kHz lie above the new Nyquist frequency of 22.05 kHz; the measurement captured what arrives at 44.1 kHz minus the tone frequency.
  • Single impulse: duration of the pre-ringing before the main peak down to −60 and −100 dB.
  • Masters: the test mix, regenerated at 48 kHz, once with a limiter at −9 LUFS and −1 dBTP, once hard-clipped at −1 dBFS to −8 LUFS. The hi-hats made from filtered noise reach up to 24 kHz; the share above 20 kHz was 28.5 and 27.3 dB below the total energy. True peak per ITU-R BS.1770-5 and a 16-times cross-check, plus sample peak and samples above 0 dBFS.
  • Processing time: the limiter master converted three times per method on a Raspberry Pi 4, counting the shortest time including reading and writing the files.

Passband: ffmpeg’s default already attenuates below 20 kHz

Method 18 kHz 19 kHz 20 kHz 20.5 kHz 21 kHz
swr, filter_size 16 −0.81 dB −1.67 dB −3.03 dB −3.95 dB −5.04 dB
swr, default (filter_size 32) −0.01 dB −0.22 dB −1.28 dB −2.43 dB −4.18 dB
swr, filter_size 128 0.00 dB 0.00 dB 0.00 dB 0.00 dB −1.01 dB
soxr, default (precision 20) 0.00 dB 0.00 dB −0.01 dB −0.07 dB −3.75 dB
soxr, precision 28 0.00 dB 0.00 dB 0.00 dB −0.03 dB −4.06 dB

With default settings, swr lost 0.22 dB at 19 kHz and 1.28 dB at 20 kHz; with filter_size 16 it was already 3.03 dB at 20 kHz. With filter_size 128 and with both soxr levels, the level stayed within 0.01 dB up to 20 kHz and within 0.07 dB up to 20.5 kHz; the steep drop only begins above that.

Above 22.05 kHz: aliasing between −9.5 and below −180 dB

Method Tone 22.5 kHz, alias at 21.6 kHz Tone 23 kHz, alias at 21.1 kHz Tone 23.5 kHz, alias at 20.6 kHz Tone 23.9 kHz, alias at 20.2 kHz
swr, filter_size 16 −9.5 dB −11.5 dB −13.8 dB −15.8 dB
swr, default (filter_size 32) −14.3 dB −19.9 dB −27.1 dB −34.4 dB
swr, filter_size 128 −96.2 dB −94.7 dB −104.0 dB −97.9 dB
soxr, default (precision 20) −136.6 dB −148.6 dB −141.6 dB −162.2 dB
soxr, precision 28 below −180 dB below −180 dB below −180 dB below −180 dB

A tone at 22.5 kHz cannot be represented at 44.1 kHz; if it remains, it appears mirrored at 21.6 kHz. The swr default attenuated it by only 14.3 dB, filter_size 16 by 9.5 dB. Only further away from the Nyquist frequency did the attenuation grow, for the default up to 34.4 dB at 23.9 kHz. swr with filter_size 128 reached 94.7 to 104.0 dB, soxr 136.6 to 162.2 dB and with precision 28 more than 180 dB. Smith names the price of steep filters: for a given stop-band specification, a lowpass filter needs approximately twice as many calculations per sample for each halving of the transition band width.

Peaks: the true peak stayed, the sample peak rose by up to 1.3 dB

Master True peak 4 times True peak 16 times Sample peak Samples above 0 dBFS
Limiter, −9 LUFS, −1 dBTP: Master at 48 kHz −1.05 dBTP −0.79 dBTP −2.31 dBFS 0
Limiter: swr, filter_size 16 −1.16 dBTP −1.07 dBTP −1.53 dBFS 0
Limiter: swr, default (filter_size 32) −1.03 dBTP −0.96 dBTP −1.44 dBFS 0
Limiter: swr, filter_size 128 −1.04 dBTP −0.98 dBTP −1.35 dBFS 0
Limiter: soxr, default (precision 20) −1.05 dBTP −0.95 dBTP −1.40 dBFS 0
Limiter: soxr, precision 28 −1.05 dBTP −0.95 dBTP −1.41 dBFS 0
Hard clipping at −1 dBFS, −8 LUFS: Master at 48 kHz +0.38 dBTP +0.72 dBTP −1.00 dBFS 0
Hard clipping at −1 dBFS: swr, filter_size 16 +0.28 dBTP +0.38 dBTP +0.01 dBFS 2
Hard clipping at −1 dBFS: swr, default (filter_size 32) +0.41 dBTP +0.54 dBTP +0.18 dBFS 18
Hard clipping at −1 dBFS: swr, filter_size 128 +0.47 dBTP +0.64 dBTP +0.27 dBFS 41
Hard clipping at −1 dBFS: soxr, default (precision 20) +0.45 dBTP +0.64 dBTP +0.26 dBFS 42
Hard clipping at −1 dBFS: soxr, precision 28 +0.44 dBTP +0.64 dBTP +0.26 dBFS 44

Loudness stayed the same for every conversion (resolution 0.1 LU). The true peak measured 4 times changed by no more than 0.11 dB; in the 16-times cross-check it fell by 0.08 to 0.34 dB, most with filter_size 16, which attenuates the highs most. The samples, by contrast, hit different points of the waveform after conversion. For the limiter master, the sample peak rose from −2.31 to as much as −1.35 dBFS. The hard-clipped master had no sample above −1.00 dBFS at 48 kHz; after conversion, 2 to 44 samples were above 0 dBFS, 42 and 44 with soxr. In a 16-bit or 24-bit file, such values are clipped. The limit on the samples at 48 kHz therefore does not protect the 44.1 kHz file, whereas the true peak of the 48 kHz file already showed beforehand how high the values can reach.

Pre-ringing and processing time: hardly any differences

Method Processing time for 32 s of stereo faster than real time Pre-ringing down to −60 dB down to −100 dB
swr, filter_size 16 0.50 s 64 times 0.14 ms 0.18 ms
swr, default (filter_size 32) 0.55 s 58 times 0.27 ms 0.36 ms
swr, filter_size 128 0.81 s 39 times 0.68 ms 1.41 ms
soxr, default (precision 20) 0.70 s 46 times 0.98 ms 1.93 ms
soxr, precision 28 0.79 s 40 times 1.32 ms 2.27 ms

All five methods rang for the same length before and after the impulse, as linear-phase filters do. Down to −60 dB, the pre-ringing lasted from 0.14 ms with filter_size 16 to 1.32 ms with soxr at precision 28, down to −100 dB no more than 2.27 ms. The more accurate methods cost little time on the Raspberry Pi 4: 32 seconds of stereo took 0.55 s with the swr default, 0.70 s with soxr and 0.79 s with precision 28, each including reading and writing the files.

What this means for the export

  • For a 44.1 kHz file from a 48 kHz export, -af aresample=44100:resampler=soxr changed the level by no more than 0.01 dB up to 20 kHz and attenuated content above 22.05 kHz by at least 137 dB; :precision=28 cost 0.1 s more.
  • -ar 44100 on its own uses the swr default: 1.3 dB less at 20 kHz and only 14 dB of attenuation for content just above 22.05 kHz. Within swr, filter_size=128 brought the deviation up to 20 kHz below 0.01 dB and the attenuation to 95 to 104 dB.
  • What matters for the peaks after conversion is the true peak of the 48 kHz file, not its sample peak. The LUFS and True Peak Meter measures both in the browser.
  • Conversion to 16-bit with dither follows as the last step after the sample rate conversion; see Dither When Exporting to 16-Bit.
  • The FL Studio manual describes the same principle for interpolation when transposing samples: linear interpolation can cause aliasing, and for the final render it recommends settings above 64-point sinc if aliasing is audible with lower methods. The setting is covered in the mastering tutorial for FL Studio.

Context and limits

The measurement covered the resamplers in ffmpeg, not the conversion in FL Studio or other programs. The test mix is synthetic; its hi-hats made from filtered noise reach up to 24 kHz, while many recordings contain less above 20 kHz. A listening test was not part of the measurement: the folded content lands between 20.2 and 21.6 kHz, and whether it becomes audible there depends on material, hearing and playback. The processing times apply to the Raspberry Pi 4 used for the measurement; the absolute values change on other hardware.

Questions and answers

Why does the more accurate resampler produce more samples above 0 dBFS on the hard-clipped master?

Because it reproduces the waveform between the samples more faithfully. Hard clipping at −1 dBFS creates edges with content reaching up to the top of the band, and the waveform described by these samples swings beyond the clipping level between them: the true peak of the 48 kHz file was +0.72 dBTP measured with 16-times oversampling, although no sample exceeded −1.00 dBFS.

A resampler calculates new samples at different points in time, so it reads this waveform at new positions and also hits spots close to the peaks. The flatter its passband, the more of the high frequencies that make up the overshoots survive. filter_size 16 already attenuates by 3.03 dB at 20 kHz, smoothing the peaks, and ended up with 2 samples above 0 dBFS; soxr and filter_size 128 leave the high end intact and ended up with 41 to 44.

The problem therefore lies in the master, not in the resampler. What matters is the true peak of the 48 kHz file: if the file is lowered until that true peak sits below 0 dBTP with a little margin, or if limiting only happens after the conversion, the samples of the 44.1 kHz file stay below 0 dBFS, at 16 or 24 bits as well.

Can the aliasing of swr also be fixed with a lower cutoff instead of a larger filter_size?

Only at the expense of the high end. The cutoff, for swr a 6 dB point at 0.97 of half the sampling rate, around 21.4 kHz, determines where the transition band lies; filter_size determines how wide it is. With the default of 32 it is so wide that it reaches both below 20 kHz and beyond 22.05 kHz: hence the 1.28 dB loss at 20 kHz and only 14.3 dB of attenuation at 22.5 kHz.

A lower cutoff merely shifts this wide band downward. The attenuation just above 22.05 kHz then rises, but the loss below 20 kHz grows along with it; this was not measured, it follows from where the transition band lies. The transition only becomes narrower with a longer filter, and that is exactly Smith’s rule of thumb: every halving of the transition band costs roughly twice as many calculations per sample. On the Raspberry Pi used for the measurement that price was small, with filter_size 128 taking 0.81 instead of 0.55 s for 32 seconds of stereo.

Does it matter whether the conversion to 16-bit comes before or after the conversion to 44.1 kHz?

Yes. The resampler calculates new samples that no longer sit on the 16-bit grid, so its output has to be rounded a second time, either without suitable dither or with a second layer of noise. On top of that come the samples above 0 dBFS that the conversion can produce and that are clipped in a 16-bit file. That is why the conversion to 16-bit with dither comes last.

Lukas Wojcik

Lukas Wojcik

Systems architect and technology enthusiast specializing in scalable tracking solutions, GMP Stack (GA4 & GTM), and robust backend architectures. Advocate for clean code and privacy-first design.

Get in Touch

Briefly describe your project or inquiry for a tailored response. This site is protected by reCAPTCHA.

Write a comment

Experience with other DAWs and plugins, differing readings and questions are welcome here.

The email address is not published. Required fields are marked with an asterisk.

ALL ARTICLES & CATEGORIES

CCTV

Follow this category by RSS

Cloud & AI

Follow this category by RSS

Data Privacy

All 16 articles in this category Follow this category by RSS

Digital Analytics

All 53 articles in this category Follow this category by RSS

Digital Marketing

All 36 articles in this category Follow this category by RSS

IT & Networks

All 17 articles in this category Follow this category by RSS

Music Production

All 14 articles in this category Follow this category by RSS

Raspberry PI

Follow this category by RSS

Smart Home

All 18 articles in this category Follow this category by RSS

Web Development

Follow this category by RSS

WordPress Plugins & Tricks

Follow this category by RSS