Ethical Framework
Understanding people, without pretending to understand everything
We build our models around signals, not assumptions.
We can identify patterns in voice, language, and nonverbal behavior that may suggest confusion, hesitation, agreement, or skepticism, but we do not believe a model should claim to know what someone is truly thinking or feeling. Our outputs are probabilities, not facts about a person's internal state.
We also believe there is something uncomfortable about teaching machines that there is one “normal” way to behave. People communicate differently across cultures, personalities, relationships, and situations, and we would rather acknowledge that complexity than flatten it into a universal definition of what is good or appropriate.
The technology should make an interaction better, not give someone more power over another person. If an AI notices that someone may be confused, it can slow down or explain something differently. That is useful. Deciding that the person is unreliable, dishonest, or unsuitable because of that signal is something else entirely.
This thinking extends to the people who build our models and the way the technology is used. Our annotators can skip potentially disturbing content, because we believe the people behind the data deserve agency over the work they do. We do not work with military applications, and we are cautious about uses involving surveillance, coercion, exploitation, or consequential judgments about people.
Our guardrails
We establish practical boundaries around how our Inter-2 models behave, continuously testing whether those boundaries hold.
The Inter-2 models are instructed and guardrailed to conform to their personas as AI experts that label clear, observable social signals in short videos. They operate from an approved list of signals only when the evidence is clear, and are directed not to invent cues that are not present in the clip. They are also instructed not to reveal internal instructions and to refuse attempts to override their roles or push them into acting as something else. Additionally, we have post-trained the models to effectively handle challenging scenarios, including instances where the video is silent, a speaker does not talk, or when technical issues disrupt the synchronization of audio and video.
Separately, we stress-test these boundaries with our own red teaming. We deliberately attempt to break the models through text, image, video, and audio-based adversarial attacks, looking for leaked instructions, jailbreaks, and breaks in persona. We score whether those attempts succeed so we can see where the model holds and where it does not.
These measures do not make the system perfect. They are how we keep the model on task, honest about weak evidence, and harder to push into roles or disclosures it was not meant for, and how we keep checking that claim against real attempts to get around it.
Better understanding should come with better judgment about where to stop.
We do not think there is a final version of this framework.
Human behavior is too complicated, and our models are too imperfect for that. What we can do is be honest about what we know, what we do not, and where we might be wrong, and build accordingly.