YAQUIZO
Create free

Audio Editor

Edit on the waveform, repair on the spectrogram, and master to the loudness target your platform actually publishes. Broadcast-standard metering, and nothing is uploaded.

Drop audio here to start editing

MP3, WAV, M4A, OGG, FLAC — or a video, to edit its soundtrack.

Nothing is uploaded. Every edit happens on this device.

Free online audio editors almost all stop at the same place: trim, fade, a volume slider, export. That covers the easy half of the job and none of the half that decides whether a recording is deliverable. This tool is built around the other half.

The spectrogram is an editing surface rather than a picture. A phone notification part-way through a sentence occupies a small patch of time and frequency; the voice under it occupies different frequencies at the same instant. Draw a box around the notification and it goes, and the sentence does not. There is no way to do that on a waveform, because on a waveform those two sounds are the same sample.

And the meters measure loudness rather than level. Peak level is nearly useless as a guide to how loud something sounds, which is why every streaming service and broadcaster now specifies LUFS. The mastering step measures your file, applies the difference, limits the true peaks, and then measures again — because limiting removes energy and therefore lowers the loudness it was applied to protect, so a single pass always lands short.

How to use it

  1. 1. Open a file

    Drop in an MP3, WAV, M4A, OGG or FLAC — or a video, to edit its soundtrack. It is decoded in your browser and never uploaded, whatever its size.

  2. 2. Select and edit on the waveform

    Drag to select, then trim, delete, silence, reverse, fade or change level. Deletes are crossfaded so edits do not click, and every change can be undone.

  3. 3. Repair on the spectrogram

    Switch to the spectrogram to see the recording as frequency over time. Arm the box tool, drag a rectangle around an unwanted sound, and only that patch is removed — the voice underneath it survives.

  4. 4. Change the voice, if you want to

    The Voice tab moves pitch and formants independently. Formants are the resonances of the throat and mouth, and they are what actually tells you how big a speaker is — which is why moving pitch alone gives you a chipmunk rather than a child. Presets range from a subtle Deeper through a middle-aged radio voice to Giant, and there are whole production chains for podcast, broadcast, audiobook and voiceover work.

  5. 5. Shape the tone

    Five bands of EQ, a compressor, a de-esser and noise reduction. Nothing is committed until you apply it, and Undo takes it straight back off.

  6. 6. Master to a delivery target

    Choose the platform you are delivering to and press master. The tool measures the loudness, applies the difference, limits true peaks to the ceiling, then measures again and corrects the error the limiting introduced.

  7. 7. Export

    WAV for anything going on to another editor, M4A for delivery, OGG Opus for the smallest file at a given quality.

Changing a voice properly

Almost every voice changer moves pitch and stops there, which is why almost every voice changer sounds like a toy. Pitch is only half of what identifies a voice. The other half is the formants: the throat and mouth act as a resonant tube that reinforces particular frequencies, and those bands stay where they are regardless of the note being spoken.

This is why a child does not sound like a child because of their pitch. They sound like a child because their vocal tract is short, which puts the formants high. Raise a man's pitch by a fifth and leave the formants alone and you do not get a child — you get the chipmunk. Move the formants without the pitch and you get the same note, spoken by a person of a different size, which is the effect no pitch shifter can produce at all.

So the two are separate controls here, and every preset moves them by different amounts — because in real bodies the two are related without being proportional. A larger person has both a longer tract and heavier folds, but the tract changes less than the pitch does. Moving them equally overshoots into caricature; moving pitch alone lands in cartoon.

The delivery targets

Destination Loudness True peak
Spotify, Amazon Music−14 LUFS−1 dBTP
YouTube−14 LUFS−1 dBTP
Apple Music, Apple Podcasts−16 LUFS−1 dBTP
Podcast (spoken word)−16 LUFS−1 dBTP
Broadcast — EBU R128−23 LUFS−1 dBTP
Broadcast — ATSC A/85−24 LUFS−2 dBTP

Delivering louder than the target does not make a track louder. The service normalises it back down on playback, and the only lasting effect is the dynamic range that was crushed to get there. Delivering quieter means sitting below everything around it.

What it is not

It is a single-file editor, not a multitrack workstation. There is no mixing several recordings together, no MIDI, no plugin hosting and no automation lanes. If you are assembling a session from many sources, use a DAW.

Spectral repair works on sounds that are separable in frequency from what is underneath them. A beep, a squeak, a click or a ring separates cleanly. A cough directly over a word shares most of its frequency range with that word, and removing it will take some of the word too — the tool will do it, but the result is a compromise rather than a rescue. And no processing removes echo, which is the original sound arriving late rather than a separate one.

Frequently asked questions

What makes this different from other free online audio editors? +

Two things, both of which are normally found only in paid desktop software. The first is spectral repair: the spectrogram is an editing surface, so you can draw a box around a single unwanted sound and remove it without touching the audio around it. The second is proper loudness metering and mastering — integrated LUFS, loudness range, and true peak measured to the ITU-R BS.1770 standard, with one-click mastering to the published targets for Spotify, YouTube, Apple, and EBU R128 broadcast. Most free online editors offer trimming and a volume slider; neither of these is in that category. Everything still runs in the browser, so nothing is uploaded.

What is spectral repair and when would I use it? +

A spectrogram shows the recording as frequency against time, so a sound occupies a patch rather than a moment. A phone notification, a chair squeak, a bird outside the window or a single cough sits in a small region of that picture, while the voice underneath occupies different frequencies at the same instant. Selecting that patch and removing it takes out the interruption and leaves the speech. Cutting on the waveform would remove the speech as well, and an EQ notch would hollow out those frequencies for the entire recording to fix one second of it.

What is LUFS, and why does it matter more than peak level? +

Peak level tells you the highest sample in a file. It says almost nothing about how loud that file sounds — a heavily compressed track and a dynamic one can both peak at exactly 0 dBFS while one is obviously much louder. LUFS measures perceived loudness instead, by filtering the signal the way the ear weights frequency and averaging over time with the quiet parts gated out. Every streaming service and every broadcaster now specifies a LUFS target, so it is the number that decides whether your file plays at the right volume.

Why does mastering louder than the target not help? +

Because the service turns it back down. Spotify, YouTube and Apple all normalise to their target on playback, so a master delivered at −8 LUFS against a −14 target is simply attenuated by six decibels. What does not come back is the dynamic range that was crushed to reach −8 in the first place. The result is a track that plays at the same volume as everyone else but sounds flatter than it needed to. Hitting the target is not a compromise; it is the whole point.

What is true peak, and why is it different from sample peak? +

Sample peak is the largest value stored in the file. True peak is the largest value the waveform actually reaches between those stored samples, once a converter or a lossy encoder reconstructs the curve through them. A file reading exactly 0.0 dBFS at every sample can still overshoot above full scale in between and distort on playback — which is why delivery specifications ask for a ceiling of −1 dBTP rather than 0. The meter here estimates true peak by reconstructing between the samples, and the limiter works to that figure rather than to sample peak.

Why does the mastering step measure twice? +

Because limiting changes the thing it was applied to protect. Applying the gain needed to reach a target and then limiting the peaks removes energy, which lowers the loudness — so a single pass lands consistently short of where it aimed. The tool measures again after limiting, corrects the residual error, and reports what it actually achieved rather than what it intended. If a tool tells you the target without measuring the result, it is telling you its intention.

How do I make a voice deeper without it sounding like a monster? +

Use the Deeper preset, or move both sliders down together. The mistake that produces a monster is moving pitch on its own: pitch is only half of what tells you whose voice you are hearing, and the other half is the formants — the fixed resonances of the throat and mouth, which do not move when you sing a lower note. Drop pitch alone and you get a slowed-down tape. Drop the formants with it, by a smaller amount, and you get a larger person. Every preset here moves the two by different amounts for that reason.

What is a formant, and why does it matter more than pitch? +

The vocal folds make a buzz at a pitch; the throat and mouth then act as a resonant tube that reinforces certain frequencies and suppresses others. Those reinforced bands are the formants, and they depend on the size and shape of the tube rather than on the note. A child does not sound like a child because their pitch is high — it is because their vocal tract is short, which puts the formants high. That is why a pitch shifter alone cannot make an adult sound like a child, and why this tool gives the two separate controls.

Does the voice change make the recording longer or shorter? +

No. Pitch is moved by a phase vocoder, which stretches the recording in time without touching the pitch, after which resampling puts the length back and leaves the pitch moved. Speed and pitch stay independent, so a lower voice speaks at exactly the same pace.

What do the production chains do? +

Each runs a whole workflow in the order an engineer would: clean the noise, tame the sibilance, shape the tone, control the dynamics, then set the delivery loudness last. Last matters — every stage before it changes the loudness, so measuring earlier would be measuring something that no longer exists. Podcast finishes at −16 LUFS, Broadcast at −23, Audiobook at −20 and Commercial voiceover at −16, each with the matching true-peak ceiling.

Is my audio uploaded anywhere? +

No. The file is decoded and every operation runs in your browser. Nothing is transmitted to any server at any point, there is no account, and closing the tab is the end of it. For interviews, medical recordings, unreleased music and confidential meetings, that is the difference between usable and not.

What does the de-esser do that an EQ cannot? +

An EQ turned down at 6 kHz removes the harshness of an "s" and takes all the air and detail with it, for the whole recording. The de-esser looks only at the sibilant band, frame by frame, and turns it down only in the frames where it is genuinely too loud. Everywhere the speaker was not hissing, the brightness is untouched.

What does tightening pauses do? +

It finds every silence longer than a second and shortens it to about a third of a second, rather than removing it. Deleting pauses outright is what makes edited speech sound breathless and wrong — the rhythm of talking is partly the gaps. Trimming them keeps the beat and loses the dead air, which is usually what you actually wanted.

Can I undo? +

Yes, up to twenty steps, on every operation including the processing ones. The history keeps whole copies of the audio rather than trying to reverse each operation, which uses more memory but cannot drift out of step with what you are looking at — and drift is the failure mode that loses work.

Is there a file size limit? +

None imposed, since nothing is uploaded. The practical limit is memory: decoding produces raw samples at roughly ten megabytes per minute of stereo CD-quality audio, and the undo history keeps copies. An hour of audio is comfortable on a normal computer; phones have far less room.

Can it replace a desktop editor? +

For a single file — editing, repairing, cleaning and mastering it to a delivery target — it does the jobs that matter and does them properly. What it is not is a multitrack workstation: there is no mixing several recordings together, no MIDI, no plugins and no automation. If you need to assemble a session from many sources, use a DAW. If you need to take one recording and make it deliverable, this does that without installing anything or uploading your audio.

Other free tools