Two schemes, and the difference between them matters more than anything else on this page. Both put an identifier into a music file and both read it back. One is audible and one is not.
The audible scheme
This is a port of WebBeeps, Danny Ayers' earlier system, which was live at webbeep.it from about 2011. It writes the identifier as music: each character becomes a pair of tones, and the tones are added to the audio.
Each character is split into two nibbles. The high nibble picks a note from a table of eight low frequencies between 261 and 440 Hz; the low nibble picks one of sixteen higher frequencies between 523 and 1760 Hz. The two are sounded together, which makes a dual tone. Tone length carries a second piece of information, so a character is read from two consecutive time slots rather than one.
Reading it back is the same idea inverted. The decoder cuts the audio into slots, measures the Goertzel power at each of the table frequencies in each slot, and reconstructs the characters from which frequencies appeared and how long they ran.
It is audible. A tone at half the amplitude of a loud passage is plainly audible. It is here because files marked with it exist and should keep reading, and because it is the baseline the other scheme is measured against.
The watermark scheme
The watermark hides the identifier in the music as a faint noise-like signal, spread across the spectrum between 250 Hz and 8 kHz. Each bit of the message is a burst of keyed noise, 2048 samples long. It is shaped, frame by frame and critical band by critical band, to sit 6 dB under the threshold that a simplified model of hearing says the music masks, so it is loud where the music is loud and nearly silent where it is not.
The noise is generated from a key. Without the key there is nothing to correlate against, so the mark cannot be found, and with it the reader recovers each bit by correlating the track against that noise, which averages away the music itself. The message is error-corrected with a convolutional code, shuffled so that damage to one stretch of the file does not land on one stretch of the message, and repeated end to end for as long as the track lasts. The reader finds where a copy starts from a keyed marker, with no help, so a cropped or shifted file reads, then combines every copy it can find before deciding. It also estimates the mark's timing from the mark itself, so a file that has been slowed or sped up by a few percent, or that has drifted by a clock error as small as thirty parts per million, is corrected and read, and the page says by how much.
Stereo files stay stereo. Every channel carries the same message, which is what lets a fold to mono, or either channel on its own, still read.
What has been measured. On eighteen full-length stereo tracks, with the mark 6 dB under the modelled masking threshold, the message read back exactly after gain changes, dither, white and pink noise at 20 dB below the loudest sample, low-pass and high-pass filtering, resampling to 48 kHz, a time shift, a crop, a fold to mono, one channel alone, MP3 compression down to 48 kbit/s, and a clock error or tempo change of up to a tenth of a percent, on every track. A change of 1% or 4% in speed, resampling to 22.05 kHz and MP3 at 32 kbit/s were lost on a few tracks each. Unmarked tracks were never reported as marked. The tables are in the project documentation. A track too short to hold one copy is refused: about 12 seconds for a few characters of identifier, 26 for 22 characters and about a minute for 63.
What has not. Nobody has yet confirmed by listening that the mark is inaudible on any particular track. The level is a measurement of signal, and a signal difference is not the same thing as a difference you can hear. Play the marked file against the original before you rely on it. And nothing has been measured against someone who is trying to remove the mark: the key stops anyone reading it, and does not by itself stop a determined person degrading it.
The marked file is written as a WAV at the same bit depth as the original, so a 24-bit master stays 24-bit. An MP3 has no bit depth of its own and becomes a 16-bit WAV.
How a reader tells "no mark" from "a mark"
This is the part that decides whether the system is safe to use. A reader handed an unmarked file must be able to say so, rather than reporting whatever it managed to read.
Every watermark payload is wrapped in a frame that begins with the four bytes
FLMK, then a version, flags, a length, and a CRC-16 of the payload. (The
audible scheme keeps its original single checksum byte, so that old files still read.) A reader
that does not find the magic has found nothing and says nothing was found. A reader that
finds the magic but whose checksum does not match says the file holds something damaged.
Those are different answers, and the page reports them differently, because "no mark here"
and "an empty identifier" are not the same thing to a person.
The checksum also catches a transposition, which a plain sum of bytes cannot: the sums of
ab and ba are equal, and the checksum of the frame is not.
The payload
The frame's payload is bytes, and what goes in them is a separate decision that has not been made yet. The intention is a header declaring that a mark is present, some metadata, then an identifying IRI, then arbitrary text, with the IRI as the part that is always present and always reliable.
What happens to your file
Nothing. The audio is decoded in your browser, the codec runs in your browser, and the result is a file your browser downloads. There is no upload and there is no server that sees your music. The page makes no request to anywhere else at all.
If you give a key, it is used in your browser to generate the noise and is never sent anywhere. The key is a phrase and is not stored. Without one the mark uses a default key that is published in the source, so anyone can find and read a mark made that way.