What makes a real-time AI avatar feel real
It is not the polygon count. Gaze, blinks, breath, timing and stillness are what make a talking face feel present. What we learned building Facelet.
We spent months making Facelet’s characters feel present, and almost none of that time went into making them prettier. Realism in a talking face is mostly about behaviour and timing. Here is what actually matters, in the order we learned it.
Gaze
A face that stares is unsettling; a face that never looks at you is absent. People hold eye contact while listening, glance away while thinking, and come back when they start to speak. Facelet’s characters do the same, with small head turns rather than eye flicks, because that is how people do it at a conversational distance.
Blinks, breath and the shoulders
A still image with a moving mouth is a puppet. Regular blinks with occasional double blinks, a chest that rises, and shoulders that shift by a millimetre or two every few seconds are what tell your eye that a body is attached to the face. Facelet animates the head, neck, breathing and shoulder sway continuously, at a level you barely notice until it is switched off.
Listening and thinking
The most convincing moments are the ones where nothing is being said. A tilt of the head while you talk, a slightly lowered gaze in the half second before an answer: these are the cues that make a reply feel considered rather than generated. Facelet has distinct listening and thinking states and moves between them as the conversation does.
The mouth, and the clock
Lip movement is the part everyone judges first, and the hardest to get right. The mouth has to open and close with the sounds, not just the volume, and it has to be on time. When the face is rendered from a voice that is already playing, the animation can only follow, and any lag reads as a dubbed film. Facelet analyses the voice as it plays, in very small slices, and drives a full facial rig from it, so the mouth shapes are real speech shapes. We say the mouth can trail the voice by a fraction of a second because it can, and we would rather you knew.
Stillness on request
Oddly, being able to stop matters as much as moving. When you ask a person to hold still, they do, and the fact that they can makes their usual motion feel chosen. Facelet’s hold-still setting freezes the idle motion; a nod or a wink on cue does the opposite. Both make the character feel like it is taking direction, because it is.
Framing and light
Facelet frames each character like a video call: face and shoulders, close, in a lit room. That distance is where skin, hair and eyes are judged, so the characters are rendered at full resolution on your Mac’s graphics chip rather than streamed as compressed video. It is also why Facelet needs an Apple silicon Mac, M2 or newer.
None of this is finished; realism never is. But it is why we describe Facelet as a face rather than an avatar, and why we would rather you watch it listen than read about it.