
Interhuman releases Inter-2 to detect social signals in real time
The AMW Read
Inter-2 adds a specialized multimodal perception model to the player map, but the report provides no independent results or adoption evidence to establish broader impact.
Interhuman releases Inter-2 to detect social signals in real time
Copenhagen-based Interhuman AI has released Inter-2, the first model in a new family designed to analyze social signals across text, audio and video in real time. The company says it detects 12 signals, including engagement, hesitation, uncertainty, confusion, agreement and disagreement, using cues such as facial expression, tone of voice, gaze and posture. Interhuman claims up to four times faster inference and improved benchmark performance, but the supplied report does not provide benchmark results or a comparison baseline.
The release places Interhuman among model builders working on how AI interprets human interaction, rather than on generating more fluent responses alone. Its proposed layer would turn nonverbal and contextual cues into structured signals that another AI system could use. That could matter in conversational products where the timing or tone of a response affects the experience, though detecting a cue is different from reliably understanding what a person means. The report describes the model's capabilities through company claims; it does not establish how well those signals generalize across people or settings.
For builders considering this capability, the practical question is whether Inter-2 improves an actual interaction under real-time conditions. Evaluations should test each signal against the intended use case, measure latency alongside accuracy, and check how often the model mistakes an ambiguous gesture or change in tone for agreement or confusion. For investors, evidence of dependable performance and adoption would say more about the opportunity than a speed claim without its underlying benchmark.