A speech model tracks the emotion of what it says, and its voice carries only part of it
Orpheus-3B was trained only to read text aloud. Inside, it keeps a running read of how the sentence it is speaking feels, and almost none of that reaches the voice.
In short
Orpheus-3B is a text-to-speech model: text goes in, a voice comes out. Nobody trained it to have opinions about the text. Yet a simple linear read-out of its internal activity recovers, frame by frame, how emotionally intense the speech it is producing is (R² 0.7306 at layer 7, against −0.2842 when the labels are shuffled), and on longer passages it recovers the emotional shape of the story itself. What the voice does is a different matter. Across held-out narrated stories, the internal read follows the story's emotional arc at a correlation of about 0.60, while the audio the model actually produces follows that same story arc at about 0.13. Push the internal direction back in by hand and you can move the voice, but only part of the way and only within a narrow window: at a quarter of the layer's own signal size, the produced sadness score rises from 0.0001 to 0.5002, and one notch further the sentence falls apart.
Push on the emotion and listen
The direction to push along is found the plainest way there is: take the model's internal activity while it reads emotional sentences, take it again while it reads flat ones, and subtract. Adding a multiple of that difference back at layer 18, while the model speaks an ordinary sentence with no feeling in it, is the whole intervention. Strength is quoted as a fraction of how large the activity at that layer already is, so 0.25 means a nudge a quarter the size of the signal it is added to.
Every position below is a real recording. Nothing is interpolated: the model was run once per emotion and strength, the audio was transcribed to count lost words, and a speech-emotion classifier scored the result. The two numbers move together, which is the point.
The steering dial
“She charged her phone and packed it into the side pocket.”
what a transcriber hears
Steering an emotionally flat sentence is the easy case. A harder question is whether the direction can talk the model out of the emotion the words themselves carry. Mostly it cannot. Sentences whose words are joyful or sad are voiced that way to begin with — 19 of 24 joyful sentences and 22 of 24 sad ones come out as the classifier's top choice for that emotion — and they resist being steered elsewhere inside the window where the words survive. Angry sentences are different: the model voices only 8 of 24 of them as angry, and those are the ones steering can redirect. Part of that apparent success is not steering at all. Unsteered angry sentences already read as sad 29% of the time (7 of 24 clips), because the model tends to voice angry text flatly or mournfully rather than angrily, so a flip rate measured against zero would credit steering with work the model was doing anyway. Joyful and sad sentences never cross-read as another emotion unsteered, in any of the four directions, so their rates need no such correction — which is itself the cleanest form of the point that words resist.
The running read
What is being pushed on? Take 17,884 frames of internal activity from 420 generated clips, hold out a fifth of them, and fit a linear read-out from the activity at one layer to how emotionally intense the audio is at that exact moment. It works: R² 0.7306 at layer 7 and 0.7436 at layer 14. Shuffle which frame gets which label and the fit is worse than predicting the average (−0.2842); shuffle the activations instead and it is the same story (−0.2784). The read is time-local, not a single verdict smeared over the clip.
On longer material you can watch it move. Below, the model narrates a held-out story about thirty seconds long. Three lines run along the same timeline: the emotional intensity the story calls for, the intensity read out of the model's internal activity, and the intensity a listener's classifier hears in the audio. Press play and watch which two lines travel together.
Watching the read move
the story asks for —read from inside —heard in the voice —
Across the 42 held-out stories the same pattern holds on average. Three correlations are involved here and they are easy to run together, so it is worth separating them. The internal read follows the story's own arc at r = 0.5684 when the emotion rises over the story and 0.5985 when it falls. The produced audio follows that same story arc at 0.1723 and 0.1310. A third pairing, the internal read against the produced audio, happens to land in the same range (0.1682 and 0.1329) while answering a different question; only the second says how much of the story's shape reaches a listener. For stories built around a turning point the voice tracks the story at 0.0004, which is nothing at all.
The internal series is also smooth in a way a noisy read-out would not be: consecutive readings correlate at 0.7652, against 0.0015 when the same readings are shuffled in time.
The gap between holding and showing
The clearest version of the gap comes from short generated stories, each labelled sentence by sentence for six emotions. A read-out at layer 8 recovers the story's joy trajectory at R² 0.5406 on held-out sentences, against −0.6142 with shuffled activations. Then compare, story by story, how well the internal read tracks the story against how well the narration does. Same stories, same measure, two very different answers.
Steering these thirty-second narrations at one matched dose confirms which way the asymmetry runs. Only sadness rises in the produced speech, by 0.1324, and it costs intelligibility (word error 0.2843). Surprise rises a little (0.0373), joy does not move (−0.0064), and fear moves the wrong way (−0.0218). The emotion the model holds internally is a usable handle for one emotion out of four.
Not a side effect of which sounds it is making
An obvious worry: a text-to-speech model's activity is dominated by which speech sound it is producing right now, and emotion correlates with speech sounds. Maybe the emotion read-out is really a phoneme read-out wearing a disguise. The test is to delete the part of the representation that carries phoneme identity — the 44-dimensional subspace spanned by the differences between the 45 speech-sound classes, 68.4% of the total variance — and see what still works.
The geometry says the same thing from a different angle. Only 0.376% of the emotion direction's length lies inside the phoneme subspace, where a random direction would put 1.414% of itself — the emotion direction actively avoids the space the phonemes occupy. Separable, though, is not the same as compact. The representation's intrinsic dimensionality is 14.26 at layer 7 before the deletion and 19.98 after: taking the phonemes out does not leave a thin emotional residue, it leaves something at least as tangled as what it started with.
Where the internal read runs out
The internal read follows gradual arcs. It does not visibly react to a designed surprise. Some of the stories were written with a turning point — a sentence where the emotional direction reverses — and the natural prediction is a larger internal jump there than elsewhere. On the training stories that is what happened (p = 0.0023). On held-out stories it did not: the turn sentences show a difference of −0.0415 against the rest, with a confidence interval from −0.171 to 0.097 and p = 0.6952, on 13 turning sentences against 186 others. With thirteen sentences this is not a demonstration that the model ignores turning points; it is a test too small to settle the question, and the training-split result makes overfitting the likelier explanation of the two.
There is also a hard boundary on what kind of emotion is being talked about here. Two things are usually meant by emotion: how worked-up something is, and whether it is pleasant or unpleasant. Only the first is decodable. In the frame-level sweep over narrated stories, intensity reaches R² 0.0938 at the best layer while pleasantness is negative at every layer tried, −0.0933 at the same one — worse than guessing the average. Everything above is about intensity. A model reading a sentence about grief and a sentence about fury may be in much the same internal state, and nothing here would tell them apart.
Taken together the picture is lopsided in a specific way. Keeping track of how the current sentence feels is apparently cheap for this model, available at several layers, separable from the phonetic work, and smooth over time. Turning that into audible prosody is expensive, and it happens for roughly one emotion. Two follow-ups would sharpen this. Holding the words fixed while varying only the requested delivery would separate representing the sentiment of the text from representing how the text should sound, which this work cannot distinguish. And reading the same direction at the audio-token decoder rather than mid-stack would say whether the bottleneck is that the plan never reaches the decoder, or that the decoder cannot act on it.
What would change this reading
- The emotion of the produced speech is scored by a classifier trained on acted human speech and pointed at synthetic speech. It was validated first: 0.8034 top-1 accuracy and 0.9543 macro AUC over six emotions on 400 acted clips, and the intensity scorer ranks acted clips at Spearman 0.7263 and gets the direction of change inside a clip right in 97.3% of 73 tested pairs. Still, a classifier is standing in for a listener throughout.
- The stories' emotional ground truth is a language model's judgement of how a story feels, sentence by sentence, not human ratings.
- The internal read-out may be riding partly on the words themselves rather than on any emotional state. That confound was not separated: nothing here distinguishes representing the sentiment of the text from representing how the text should sound.
- The held-out arc correlations were not separated from simple drift over the course of a story, which would produce a similar number for a different reason.
- Internal fit quality and produced-speech agreement are different measurements on different objects. Where they are compared here, it is like against like: the same per-story correlation, computed for the internal read and for the audio.
- One model, one voice, one checkpoint, one set of generated evaluation data.