A speech model tracks the emotion of what it says, and its voice carries only part of it

Orpheus-3B was trained only to read text aloud. Inside, it keeps a running read of how the sentence it is speaking feels, and almost none of that reaches the voice.

In short

Orpheus-3B is a text-to-speech model: text goes in, a voice comes out. Nobody trained it to have opinions about the text. Yet a simple linear read-out of its internal activity recovers, frame by frame, how emotionally intense the speech it is producing is (R² 0.7306 at layer 7, against −0.2842 when the labels are shuffled), and on longer passages it recovers the emotional shape of the story itself. What the voice does is a different matter. Across held-out narrated stories, the internal read follows the story's emotional arc at a correlation of about 0.60, while the audio the model actually produces follows that same story arc at about 0.13. Push the internal direction back in by hand and you can move the voice, but only part of the way and only within a narrow window: at a quarter of the layer's own signal size, the produced sadness score rises from 0.0001 to 0.5002, and one notch further the sentence falls apart.

Push on the emotion and listen

The direction to push along is found the plainest way there is: take the model's internal activity while it reads emotional sentences, take it again while it reads flat ones, and subtract. Adding a multiple of that difference back at layer 18, while the model speaks an ordinary sentence with no feeling in it, is the whole intervention. Strength is quoted as a fraction of how large the activity at that layer already is, so 0.25 means a nudge a quarter the size of the signal it is added to.

Every position below is a real recording. Nothing is interpolated: the model was run once per emotion and strength, the audio was transcribed to count lost words, and a speech-emotion classifier scored the result. The two numbers move together, which is the point.

The steering dial

“She charged her phone and packed it into the side pocket.”

steer toward
off0.050.100.150.250.35
sadness heard in the voice0.000
words lost, per word spoken0.000

what a transcriber hears

One neutral sentence, sixteen real recordings. This sentence was picked before listening, as the one whose response to the sadness direction at 0.25 strength sits closest to the average across all such sentences; the curve below gives the averages. Move the dial past 0.25 and the emotion score keeps climbing while the sentence stops being the sentence. The clips are re-encoded to a compact format so the page carries them.
toward sadnesstoward joytoward angerrandom direction, same size 0.00.20.40.6emotion heard in the voiceoff0.050.100.150.250.35words break uppast heresadnessjoyangerrandom direction 0123words lost per word spokenoff0.050.100.150.250.35steering strength, as a fraction of the layer's own signal sizeone word in fivesadnessjoyanger
Averages over 40 recorded clips per point, 1,632 recorded clips across the whole sweep. The three coloured lines are the emotion the classifier hears in the produced speech when the model is steered toward that emotion; the grey line is a random direction of the same length, which never gets above 0.0096. The lower panel is how much of the sentence survives: unsteered speech transcribes at 0.0214 errors per word, sadness at 0.25 strength at 0.2132, and at 0.35 the output stops being speech at all (2.76). Sadness is the direction that works; anger barely moves inside the usable range.

Steering an emotionally flat sentence is the easy case. A harder question is whether the direction can talk the model out of the emotion the words themselves carry. Mostly it cannot. Sentences whose words are joyful or sad are voiced that way to begin with — 19 of 24 joyful sentences and 22 of 24 sad ones come out as the classifier's top choice for that emotion — and they resist being steered elsewhere inside the window where the words survive. Angry sentences are different: the model voices only 8 of 24 of them as angry, and those are the ones steering can redirect. Part of that apparent success is not steering at all. Unsteered angry sentences already read as sad 29% of the time (7 of 24 clips), because the model tends to voice angry text flatly or mournfully rather than angrily, so a flip rate measured against zero would credit steering with work the model was doing anyway. Joyful and sad sentences never cross-read as another emotion unsteered, in any of the four directions, so their rates need no such correction — which is itself the cleanest form of the point that words resist.

already reads that way with no steeringadded by steering, sentence intact 0%20%40%60%80%clips whose loudest emotion is the target+21 ptsanger wordssteered to sadnessunsteered 29%+25 ptsanger wordssteered to joyunsteered 4%+12 ptsjoy wordssteered to sadnessunsteered 0%no gainsadness wordssteered to joyunsteered 0%+4 ptsjoy wordssteered to angerunsteered 0%no gainsadness wordssteered to angerunsteered 0%
At 0.25 strength, how often a sentence's loudest emotion is the steering target while the words stay intelligible, split into what the unsteered model already did (grey) and what steering added (blue); the pale extension is extra flips that cost the sentence. Angry words steered toward sadness reach 50%, but 29% of them (7 of 24) already read as sad unsteered, so steering adds 21 points. Toward joy the baseline is 4% (1 of 24) and steering adds 25 points. Every other baseline here is 0%, so those rates are all steering: 12 points for joyful words pushed toward sadness, 4 for joyful words pushed toward anger, and nothing at all for sad words in either direction. Joyful words pushed toward sadness do flip 62% of the time if intelligibility is ignored, which is the pale bar.

The running read

What is being pushed on? Take 17,884 frames of internal activity from 420 generated clips, hold out a fifth of them, and fit a linear read-out from the activity at one layer to how emotionally intense the audio is at that exact moment. It works: R² 0.7306 at layer 7 and 0.7436 at layer 14. Shuffle which frame gets which label and the fit is worse than predicting the average (−0.2842); shuffle the activations instead and it is the same story (−0.2784). The read is time-local, not a single verdict smeared over the clip.

On longer material you can watch it move. Below, the model narrates a held-out story about thirty seconds long. Three lines run along the same timeline: the emotional intensity the story calls for, the intensity read out of the model's internal activity, and the intensity a listener's classifier hears in the audio. Press play and watch which two lines travel together.

Watching the read move

story

the story asks for read from inside heard in the voice

The vertical rules mark sentence boundaries. Both clips are held-out stories, chosen for clean transcription; the first is the one whose internal agreement with the story sits nearest the average for its arc shape. In it the internal read follows the story at r = 0.4586 while the audio manages r = 0.1072. The second clip is a stronger case, r = 0.8778 inside against r = 0.3435 in the voice.

Across the 42 held-out stories the same pattern holds on average. Three correlations are involved here and they are easy to run together, so it is worth separating them. The internal read follows the story's own arc at r = 0.5684 when the emotion rises over the story and 0.5985 when it falls. The produced audio follows that same story arc at 0.1723 and 0.1310. A third pairing, the internal read against the produced audio, happens to land in the same range (0.1682 and 0.1329) while answering a different question; only the second says how much of the story's shape reaches a listener. For stories built around a turning point the voice tracks the story at 0.0004, which is nothing at all.

internal read against the storyproduced audio against the storyinternal read against the produced audio 0.00.20.40.6agreement between the two series (r)0.5680.1720.168emotion risesn=150.5980.1310.133emotion fallsn=140.2670.0000.106turning pointn=13
Held-out stories only, averaged per story. The middle bar is the one that answers how much of the story reaches the voice; the right-hand bar is a different pairing that sits at a similar height, which is how the two get confused. Only the first and third were published by the source experiment; the middle bar is computed from its own per-story series, and that recomputation reproduces all 18 of the published aggregate values first, as a check that the pairings line up.

The internal series is also smooth in a way a noisy read-out would not be: consecutive readings correlate at 0.7652, against 0.0015 when the same readings are shuffled in time.

The gap between holding and showing

The clearest version of the gap comes from short generated stories, each labelled sentence by sentence for six emotions. A read-out at layer 8 recovers the story's joy trajectory at R² 0.5406 on held-out sentences, against −0.6142 with shuffled activations. Then compare, story by story, how well the internal read tracks the story against how well the narration does. Same stories, same measure, two very different answers.

read from inside the modelheard in the voice 0.00.20.40.60.8agreement with the story's emotion (r)joy0.710.13n=107surprise0.540.06n=105sadness0.520.17n=94fear0.51-0.01n=86disgust0.420.05n=28anger0.40-0.02n=35
Per-story agreement with the story's own emotion labels, on held-out stories, for the internal read-out and for the produced narration. Joy is the strongest case in both columns and the gap is widest there: 0.7063 inside against 0.1333 in the audio. Sadness is the best the voice manages, 0.1688. For anger the narration does not follow the story at all (−0.0219). Sample sizes differ because only stories with variation in that emotion are counted.

Steering these thirty-second narrations at one matched dose confirms which way the asymmetry runs. Only sadness rises in the produced speech, by 0.1324, and it costs intelligibility (word error 0.2843). Surprise rises a little (0.0373), joy does not move (−0.0064), and fear moves the wrong way (−0.0218). The emotion the model holds internally is a usable handle for one emotion out of four.

Not a side effect of which sounds it is making

An obvious worry: a text-to-speech model's activity is dominated by which speech sound it is producing right now, and emotion correlates with speech sounds. Maybe the emotion read-out is really a phoneme read-out wearing a disguise. The test is to delete the part of the representation that carries phoneme identity — the 44-dimensional subspace spanned by the differences between the 45 speech-sound classes, 68.4% of the total variance — and see what still works.

0.00.20.40.60.8how well it can be read out0.6880.003beforeafterwhich speech sound0.7310.677beforeafteremotional intensitychance for 45 sounds
After the deletion, identifying the current speech sound is no longer possible at all: macro F1 falls from 0.6882 to 0.0028 and accuracy to 0.0665, which is chance for 45 classes. The emotional-intensity read-out keeps 92.7% of what it had, R² 0.6773 down from 0.7306. Deleting a random subspace of the same size leaves it at 0.7299, so the small loss is real but slight.

The geometry says the same thing from a different angle. Only 0.376% of the emotion direction's length lies inside the phoneme subspace, where a random direction would put 1.414% of itself — the emotion direction actively avoids the space the phonemes occupy. Separable, though, is not the same as compact. The representation's intrinsic dimensionality is 14.26 at layer 7 before the deletion and 19.98 after: taking the phonemes out does not leave a thin emotional residue, it leaves something at least as tangled as what it started with.

Where the internal read runs out

The internal read follows gradual arcs. It does not visibly react to a designed surprise. Some of the stories were written with a turning point — a sentence where the emotional direction reverses — and the natural prediction is a larger internal jump there than elsewhere. On the training stories that is what happened (p = 0.0023). On held-out stories it did not: the turn sentences show a difference of −0.0415 against the rest, with a confidence interval from −0.171 to 0.097 and p = 0.6952, on 13 turning sentences against 186 others. With thirteen sentences this is not a demonstration that the model ignores turning points; it is a test too small to settle the question, and the training-split result makes overfitting the likelier explanation of the two.

There is also a hard boundary on what kind of emotion is being talked about here. Two things are usually meant by emotion: how worked-up something is, and whether it is pleasant or unpleasant. Only the first is decodable. In the frame-level sweep over narrated stories, intensity reaches R² 0.0938 at the best layer while pleasantness is negative at every layer tried, −0.0933 at the same one — worse than guessing the average. Everything above is about intensity. A model reading a sentence about grief and a sentence about fury may be in much the same internal state, and nothing here would tell them apart.

Taken together the picture is lopsided in a specific way. Keeping track of how the current sentence feels is apparently cheap for this model, available at several layers, separable from the phonetic work, and smooth over time. Turning that into audible prosody is expensive, and it happens for roughly one emotion. Two follow-ups would sharpen this. Holding the words fixed while varying only the requested delivery would separate representing the sentiment of the text from representing how the text should sound, which this work cannot distinguish. And reading the same direction at the audio-token decoder rather than mid-stack would say whether the bottleneck is that the plan never reaches the decoder, or that the decoder cannot act on it.

What would change this reading