inter-2-audioResearch

Inter-2 Audio: Learning Social Signals from Speech

How we trained an Audio LLM to read confidence, hesitation and uncertainty from speech alone, and to stay neutral when the evidence isn't there.

·8 min read
Inter-2 Audio: Learning Social Signals from Speech

Voice contains far more information than the words being spoken. The way someone hesitates, where they place emphasis, how their pace changes or whether their delivery becomes more certain can all provide additional context for understanding what is happening in a conversation.

As Voice AI becomes part of more human conversations, from AI-moderated research interviews and sales roleplays to voice agents, coaching tools and conversational assistants, understanding this layer of information becomes increasingly useful. A transcript can tell a system what someone said, but speech itself also contains intonation, loudness, rhythm, emphasis, speaking rate and pauses, which together can provide evidence of signals such as confidence, hesitation, uncertainty or frustration.

Inter-2 Audio is an audio model built to make this information available to AI directly from speech. Rather than treating social-signal detection as a matter of identifying individual acoustic cues, we trained the model to jointly interpret what someone says and how they say it, while dealing with a particularly important problem for voice-based systems: how do you know when an acoustic cue is actually meaningful?

Learning from what is said and how it is said

Audio models already perform well across a broad range of speech-understanding tasks, but interpreting social signals introduces a different problem. Models need to consider lexical information and acoustic information simultaneously, particularly when the two do not point toward the same interpretation.

A filled pause such as “uh”, for example, may contribute to hesitation, but it does not necessarily mean that the speaker is hesitant. A change in pitch may carry social information in one context and very little in another. Someone can speak confidently while pausing to formulate a thought, just as uncertainty can be communicated through otherwise fluent speech.

Many general-purpose audio models can over-rely on the transcript or treat particularly salient words and disfluencies as direct triggers for a prediction. Inter-2 Audio was trained to interpret these cues in relation to one another, combining lexical evidence such as verbal hedges and filled pauses with acoustic evidence including pitch movement, loudness, emphasis, speaking rate and timing.

That changes the task from identifying individual cues to evaluating the evidence across the utterance. A pause matters, but so does what surrounds it. A change in pitch matters, but so do the speaker's rhythm, loudness and overall delivery.

This complex interconnection of cues leads to another problem: sometimes none of the cues provide enough evidence to make a prediction at all.

Teaching the model when not to predict

Social-signal detection can easily become a problem of overinterpretation. If a model has been trained primarily on examples containing recognisable social signals, it can learn to search for the closest available label even when the evidence is ambiguous.

For Inter-2 Audio, we therefore trained the model to recognize neutrality explicitly, rather than treating it as the default when no social signal is predicted strongly enough..

We trained the model through a three-stage curriculum learning procedure. In the first two stages, the model learns the acoustic and lexical patterns associated with different social signals. In the third it learns ambiguity. More specifically, the stages consist of:

  • Encoder Awakening: We add and warm up adapters at selected layers of the audio encoder, allowing it to acquire target-domain acoustic knowledge while preserving its pretrained representations (Shi et al., 2026).
  • Audio Grounding: We apply supervised fine-tuning (SFT) on a curated base mixture of naturalistic interaction audio. At this stage, the model is grounded in the target task using only examples with actual social signals. The incentive is to teach the model a speech-based social-signal taxonomy, evidence-based rationales, and engagement and confidence fields.
  • Evidence Calibration: In the third stage, we perform a full dataset pass, introducing ambiguous clips that expert annotators have judged to be explicitly neutral. This teaches the model not only what a social signal can sound like, but also when the available evidence is insufficient to justify one. A hinge-margin loss is applied at the signal-presence decision, encouraging a clear separation between clips that contain a social signal and those that should remain neutral.

The curriculum learning proved to be highly beneficial as social signals frequently share acoustic characteristics. Confidence, hesitation and uncertainty, for example, can overlap in their prosodic expression, while speech that sounds confident does not necessarily mean that Confidence is the socially relevant signal. A system that simply maps individual cues to labels risks turning ambiguity into certainty.

The behaviour that emerged from this training is visible in the model's errors. When Inter-2 Audio misses a true Confidence label, its most frequent alternative prediction is neutral. The evaluated baseline models, by comparison, most frequently substitute Hesitation or Uncertainty.

Both outcomes reduce recall, but they are meaningfully different errors: one withholds a prediction when the evidence is insufficient, while the other introduces a social interpretation that may not be supported by the clip.

More information from the whole utterance

To illustrate Inter2-Audio’s performance, we wanted to understand whether the model was simply detecting more cues or whether it was using those cues differently.

One example from our evaluation shows this distinction. In a five-second clip, a speaker uses several filled pauses and a false start, both of which provide reasonable evidence for Hesitation and Uncertainty. At the same time, the speaker maintains a loud, steady delivery throughout the clip, providing evidence for Confidence. Human annotators identified all three signals as co-occurring.

Both Inter-2 Audio and Gemini 3.1 Pro recognised the local disfluencies. Gemini used those cues to identify Hesitation but missed Confidence. Inter-2 Audio also acknowledged the pauses, but considered them alongside the sustained loudness, steadiness and absence of verbal hedges across the utterance, allowing it to recover Confidence as an additional signal.

A single example cannot tell us what is happening internally within either model, but it does illustrate the behaviour we were trying to encourage during training: local cues should be interpreted in the context of the speaker's overall delivery rather than treated as isolated evidence. Across the broader evaluation, this appears alongside the model's tendency to remain neutral when the combined evidence is not sufficiently strong.

Benchmarking Inter-2 Audio

We evaluated the social-signal prediction capabilities of Inter-2 Audio on a test set of 839 audio clips, approximately 28% of which were neutral. We used a six-signal vocabulary alongside a separate three-class task measuring whether the speaker was engaged, neutral or disengaged.

Because the model is designed around signals that can be meaningfully inferred from speech, the Inter-2 Audio catalogue focuses on six audio-grounded social signals, excluding signals such as Skepticism, Disagreement and Confusion where reliable interpretation depends more heavily on visual evidence (Bousmalis et al., 2009).

We compared the model against a combination of open-weight audio models and closed frontier models served through APIs, including Qwen2-Audio, Audio-Flamingo-3, Kimi-Audio, Ultravox, GPT-Audio, Gemini 3.1 Pro, Gemini 3.8 Flash and Qwen3.5-Omni-Plus. All models received the same task instructions, social-signal catalogue and label vocabulary.

Across this evaluation, Inter-2 Audio ranked first on every reported metric, reaching:

54.67 Micro-F1 · 58.08 EW-F1 · 47.69 Neutral-F1 · 82.72% Engagement Accuracy

Compared with Gemini 3.1 Pro, the strongest baseline across the main social-signal metrics, Inter-2 Audio improves Micro-F1 by 3.47 percentage points, and Evidence-Weighted F1 by 4.10 percentage points. This is achieved at a fraction of Gemini 3.1 Pro’s inference speed.

The clearest difference, however, is in neutral detection. Inter2-Audio achieves a Neutral-F1 of 47.69, compared with 41.61 for the strongest baseline on this metric. The gap is even larger for most API-hosted foundation models, with Gemini 3.1 Pro reaching 34.17, GPT-Audio 29.47, and Gemini 3.8 Flash just 2.15. This suggests that Inter2-Audio is not simply predicting more social signals, but is better at distinguishing meaningful social evidence from ambiguous or neutral speech.

For more information about the metric calculations, please see the blog post about our Inter-2 model.

Beyond clean audio

Real conversations rarely happen under benchmark conditions. Microphones sit at different distances, rooms introduce reverberation, and background conversations, movement and environmental noise all affect the signal available to a model.

We therefore evaluated Inter-2 Audio under progressively more difficult acoustic conditions by introducing real room impulse responses and mixtures of continuous and transient noise following the FFASR evaluation procedure (Treble Technologies, 2026). The evaluation ranged from high-SNR conditions, where degradation was relatively mild, through moderate background activity and finally low-SNR conditions designed to represent noisy or distant capture.

Inter-2 Audio remained comparatively robust as those conditions became more difficult, with substantially less degradation than models such as GPT-Audio and Kimi-Audio across the tested noise conditions. This matters for applications outside controlled recording environments, where the social information contained in someone's voice needs to remain accessible even when the audio itself is imperfect.

Another layer of information for Voice AI

Speech carries information at several levels simultaneously. There are the words themselves, but also the rhythm, emphasis, pauses, intonation and changes in delivery that people naturally use when interpreting one another.

Inter-2 Audio makes more of that information available to AI by jointly interpreting what someone says and how they say it, while remaining deliberately cautious about the point at which acoustic evidence becomes a social interpretation. Across our evaluation, that approach improves social-signal detection, produces substantially stronger neutral detection and remains robust as the acoustic environment becomes more difficult.

For Voice AI, this adds another layer of information to work with, opening the possibility for products built around conversation to understand not only what was said, but more of how it was communicated.

References

  • Shi, M., Wang, Z., Shankar, N. B., Zhang, K., Eren, E., & Alwan, A. (2026). Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs.
  • Bousmalis, K., Mehu, M., & Pantic, M. (2009). Spotting Agreement and Disagreement: A Survey of Nonverbal Audiovisual Cues and Tools. IEEE Journal of Solid-state Circuits.
  • Treble Technologies (2026). FFASR Evaluation Leaderboard.