16.4 C
London
HomeTechInterhuman AI wants to teach AI what humans say without words

Interhuman AI wants to teach AI what humans say without words

Copenhagen-based Interhuman is developing proprietary models that interpret the social signals behind human communication, from tone of voice to body language.

AI has become good at understanding what we say. But human communication has never been just about words. We pick up on hesitation, confidence, uncertainty, engagement, and agreement, often without anyone overtly expressing them.

These non-verbal and contextual signals shape how we interpret an interaction, yet much of that information remains largely invisible to today’s AI systems. Copenhagen-based Interhuman AI is working to close that gap.

One of a relatively small number of European AI startups building its own models, the company develops Social Intelligence technology designed to help AI interpret these signals. Its latest model, Inter-2, analyses social signals across text, audio, and video in real time.

From ChatGPT to social intelligence

I spoke with co-founders COO Frederik Sally and CEO Paula Petcu to learn more. Petcu said the idea for Interhuman emerged with the mainstream arrival of LLMs such as ChatGPT.

While they transformed how people interacted with AI, they remained largely blind to body language, tone of voice, facial expressions, and other social signals. At the time, Petcu was working in the pharmaceutical industry, exploring how AI could better understand patients in clinical trials.

“That’s where the idea came from: could we combine large language models with models analysing video and audio from a camera and microphone, and add another layer of sensing to understand what the user is actually communicating?”

Petcu met co-founder and CRO Line Clemmense, who had been researching the field, through a webinar. The pair built an initial prototype: an AI coach that responded not only to what someone said, but how they said it.

They later joined Antler’s matchmaking programme, where they met Sally, who brought an operations and product perspective to the team.

“The user experience of AI depends enormously on how the model reacts to you and how well it understands you,” said Sally.

Inter-2 is the first model in a new family. It analyses facial expressions, tone of voice, body language, and behavioural cues such as shifts in gaze, posture, and head gestures to detect 12 social signals in real time, including engagement, hesitation, uncertainty, confusion, agreement, and disagreement. The company says Inter-2 also delivers up to four times faster inference alongside improved benchmark performance, making real-time signal detection more practical at scale. ​

“Today’s models are remarkably capable and still socially deaf. They respond to the words, not to the person who wrote them. A correct answer delivered at the wrong moment is still a bad answer,” shared Sally. ​

Working towards artificial social intelligence

Inter-2 is a step towards what Interhuman calls Artificial Social Intelligence: giving AI systems more of the perceptual ability humans use to navigate interactions every day. Interhuman combines behavioural science with machine learning, translating how people perceive and interpret one another into structured signals that AI systems can process. The company is building its own models.

According to Petcu, major AI labs largely focus on general intelligence and productivity, leaving less attention for human communication and social-signal detection.

“If you don’t account for this layer in future AI systems, there’s a risk of over-optimising workflows and productivity without considering the potential consequences.”

She points to AI-powered healthcare triage as one example. If someone is frustrated or distressed while speaking to an AI over the phone, a system that cannot perceive those signals could misinterpret the situation and make the wrong decision about emergency care. In robotics, if a humanoid robot can’t perceive that it should stop an action, it could cause physical harm.

The problem of interpreting human behaviour

But human behaviour is messy and highly individual. Personal idiosyncrasies mean a social cue can signal different things depending on the person or context. A pause or hesitation, for example, could indicate uncertainty, but it could just as easily mean someone is thinking before they speak.

Crucially, Interhuman is designing its models so that their assessments are traceable, allowing end users to understand which cues led to a particular conclusion.

Sally explained:

“We thought it was important to include the rationale behind the model’s outputs. You should be able to trace what caused a particular signal to be identified.

We’re trying to make the model as transparent as possible so you can understand what was behind an assessment.”

Imagine an AI-assisted interview, for example — you should be able to trace what happened in that conversation.

“Why did the model identify hesitation at that particular moment? As a human being, if an AI is determining something about my future, I want to be able to question it.”

Accounting for differences in culture, language, age, neurodiversity, and individual communication styles presents another challenge.

For Sally, addressing that variation starts with collecting enough diverse data and annotating it appropriately. The company initially used publicly available data to understand the problem and develop its annotation process.

It now also collects its own data, with employees recording different types of conversations, and works with external data providers to expand its datasets. Interhuman works with expert annotators, including psychologists and behavioural scientists, to label the behavioural data used to develop its models.

Sally said building this data pipeline — and training people to annotate the material consistently — represents a significant part of the company’s work.

“Our approach is grounded in behavioural-science literature, and we’re trying to make it as rigorous as possible from the beginning.”

See also
Proxima Fusion partnership gives Europe its most credible path to commercial fusion

The responsibility of building AI

As AI capabilities grow, issues around privacy, manipulation, and human control become even more important.

“Social Intelligence should not be about judging people or claiming to know what someone is thinking. It should give AI more context to understand human communication, while keeping people in control of how that information is used,” Petcu shared.

I was curious about how Interhuman ensures companies using its models do so responsibly. Sally admits its something the company discusses frequently.

Interhuman has already drawn some boundaries around who it will work with, including turning down an investor interested in defence applications for the technology. He points to AI-enabled smart glasses as an example of how social intelligence could potentially be misused:

“Imagine adding our technology to that and using it to analyse strangers as you move through the real world — potentially even using those insights to manipulate people or gain an advantage over them.”

As more companies begin using its technology, Interhuman plans to develop a broader set of rules governing acceptable uses of its models, alongside restrictions already imposed on certain AI applications under European regulation.

Petcu said the company is currently small enough to maintain visibility into how customers are using its systems and engage with them directly.

“That lets us learn more about their business, answer questions about the EU AI Act or data privacy, and potentially tell them, ‘You can’t use this technology for the purpose you’re intending.’”

Interhuman is GDPR-compliant and currently undergoing security certification audits. Its process transparency helps when talking to companies and their legal departments.

Where social intelligence could be used

Interhuman is sector-agnostic as it provides an API that gives companies access to its model. For example, it has customers building AI-based role-play applications for corporate training. That might involve practising a difficult conversation, a salary negotiation, or an interview with an AI. The AI can provide feedback not only on what someone says, but how they say it.

“That makes the feedback much more actionable because it can also assess how you present yourself,” asserts Sally.

The company is also working in digital health, including with clinical psychologists in Denmark at the Centre for Mental Health for Children and Adolescents, supporting an app that helps parents learn to communicate more effectively with their children.

Petcu detailed:

“For example, if a child doesn’t want to go to school and is screaming, how do you handle that situation? How do you make sure you’re using the right body language, tone of voice, and words?”

There’s also digital-health coaching: an empathetic AI coach that can hear how you’re saying something and potentially ask whether there’s something more behind what you’ve said.

“Then there are applications in sales training and sales intelligence, including meeting note-takers and sales coaching.”

Market research is another area.

“You might have an AI interviewer asking a customer what they think about a product’s taste. If they say it’s sweet, the AI can recognise signals around that response and probe further: is it too sweet?” explained Petcu.

In the future, the company expects social intelligence to become a layer for robotics, elderly care, AI companions, and other physical applications. It’s also talking to companies aiming to humanise their avatars — one of the problems they’re working on is something as simple as how an avatar should behave while it’s listening.

“That’s actually a difficult problem. How should it react while another person is speaking? Companies are approaching us about using our technology for that,” said Sally.

Making AI socially aware in more complex conversations

While Interhuman has focused a lot on sales training, coaching, and intelligence, many sales calls and customer-support interactions happen over the phone.

“That’s voice-only, so you don’t have visual information,” explained Sally.

Interhuman is developing a model that can handle this: “It’s obviously easier to work with a rich dataset containing both video and audio,” explained Sally.

“When you have only the voice, understanding what is happening becomes more difficult.”

But it also enables many more voice-AI applications. He predicts voice will become the next major interface for interacting with AI.

“We’ve spent a lot of time interacting through chat, but we’re already seeing developers speaking to their AI coding agents.”

From here, the company is working to improve individual signals and make the model more robust with noisy data. It also wants to support conversations involving multiple people.

“What is the relationship between those two people? How should the AI approach person A versus person B? What do they know about each other?” shared Petcu.

“There’s a lot of nuance that we still need to teach the model.”

Power dynamics are another area being explored. A job interview is very different from a friendly conversation, and the same signal can mean different things depending on the situation. Interhuman will announce two additional models in October. ​

Building AI models in Europe

For Petcu, it’s important that Interhuman is developing its own AI models in Europe:

“Building an AI lab and developing your own models that could have an impact on society is quite unusual here.”

She argues that developing the technology within Europe’s more demanding regulatory environment could ultimately give Interhuman an advantage as it expands into other markets.

“If we can build this technology successfully here and meet those requirements, I think we’ll be well positioned to meet regulatory requirements elsewhere as well.”

Lead image: Interhuman AI co-founders: CEO Paula Petcu, COO Frederik Sally, and CRO Line Clemmensen.

Latest news
Related News