Keel
Suno-class generate. Not Suno.
120 credits
Lock is a product surface

Rewrite the line. Keep the band.

Keel is an original studio in this repo. It is not Suno, does not call Suno’s API, and does not ship unlicensed celebrity voices. Duration is a hard budget. Famous names are refused.

Solved enough
Keep the music
Stem split (Demucs / UVR) or start from stems. Reverb bleed, doubles, and vocal chops in the beat still leak. The instrumental you “keep” is often a hole with ghosts in it.
Research, 2026
New words, same clock
Text-based singing voice editing. MeloDISinger (KAIST, Jun 2026) treats this as a fixed-budget duration problem: new phonemes must sum to the old span so the band does not drift. It is SOTA, not a product, and it assumes a feasible lyric.
Easy timbre, hard identity
Someone else’s throat
RVC / Seed-VC / Kits convert existing singing to another voice without changing lyrics. Synthesizer V and ACE Studio synthesize new lyrics, but against a score and a licensed voicebank. Cloning a famous singer on a hit is a likeness claim, not an ML leftover.

Pack a lyric onto a frozen melody

Same six notes, 6.5 beats. The music does not move. Your words have to.

Locked6 syllables · 15 phones · 6 notes · budget 6.5 beats

Syllable count roughly matches the melody. This is the only case where “perfect timing” and “new lyrics” can both be true without chewing the vowels.

C4 · 4/16
ay
was “I
E4 · 2/16
nehv
was “nev
G4 · 6/16
er
was “er
A4 · 4/16
lernd
was “learned
G4 · 2/16
tuw
was “to
E4 · 8/16
rehst
was “rest

These are three different models

Voice conversion keeps the words. Singing synthesis needs a piano roll. Lyric editing must not move the mix. Glue them and each one corrupts the others: convert after an edit and the new consonants smear; synthesize a cover and the drums shift; inpaint a line and the room tone dies.

Lyrics are not a subtitle track

Speech TTS can stretch a sentence. Singing cannot. Vowels are the notes. Consonants are the onsets. “Stay” and “put your name in the chorus tonight” are not interchangeable on the same MIDI. The 2026 papers even generate evaluation lyrics with an LLM so the new line is duration-feasible. Arbitrary fan lyrics fail that test.

Humans are not on the grid

“Perfect timing” usually means sample-lock to the original performance, which already leans, scoops, and breathes off the click. Forced aligners trained on speech miss long vowels. Score-based singers lock to MIDI and then sound like MIDI. You pick one clock. You do not get both for free.

Licensed singers, lyrics on a piano roll, sample-accurate to the score you wrote.
Split a song, extract a pitch/timing map, resynthesize with their voices.
This recording, that timbre. Timing and lyrics stay put.
Regenerate a section or restyle a song. Fast. Sounds like a record.
Research SOTA for duration-preserving lyric edits via duration ratios + infilling.

Pick a song in the library, then press Play.

0:000:00