Two things get called an AI avatar
One is a video. You type a script, a service renders a talking head, and a minute later you have a clip. It looks good and it cannot hear you. The other is a character that is alive while you talk to it: it is drawn sixty times a second, it notices when you start speaking, and it answers with its face as well as its voice. Facelet is the second kind.
What happens while you talk
Facelet renders a lifelike 3D human on your Mac and keeps it in one of three states. Listening: the character looks at you, blinks, breathes, and holds still enough that you can tell it is paying attention. Thinking: ChatGPT is composing its answer, and the character glances away the way a person does when they are choosing words. Speaking: the answer arrives as ChatGPT’s own voice, and Facelet moves the mouth, jaw, brows and eyes from that audio as it plays.
None of this is a recording. Every frame is drawn on your Mac in the moment, so the face reacts as the words arrive rather than after a render finishes. That is what real time means here, and it is the difference between watching a face and being with one.
A human, not a cartoon
The ten characters in the first collection are lifelike humans, with skin, hair and eyes rendered in 3D at a quality you would expect from a modern game or film pipeline. They are not anime-style models, not a photo with a moving mouth, and not a webcam filter. Each has a resting expression of their own, and each remembers which of ChatGPT’s voices you gave them.
Honest about the mouth
Facelet animates the mouth from the audio it is playing, on your Mac. Because the voice is being generated live by ChatGPT, the mouth can trail the sound by a fraction of a second. We would rather tell you that than promise studio lip sync. In practice the timing reads as natural because the rest of the face, the gaze, the blinks and the breath, is doing most of the work. We wrote about that in what makes an AI avatar feel real.
Gestures on cue
Ask for a nod, a shake of the head or a wink from the Gestures menu, or tell the character to hold still. Everything else happens on its own. The aim is a face you stop noticing, the way you stop noticing a person’s face once you are talking to them.
What it needs
Rendering a human at a smooth frame rate while a voice conversation runs takes a modern graphics chip, so Facelet runs on Apple silicon Macs with an M2 chip or newer and macOS 14 or newer, with 16 GB of memory recommended. There is no browser version, because a browser cannot do this well. There is also nothing to wait for and nothing billed by the minute: the rendering is yours, on your machine. See rendered on your Mac.
Questions
What does “real time” mean for an AI avatar?
The face is drawn frame by frame while the conversation happens, rather than generated as a video clip after the fact. Facelet renders its character live on your Mac’s graphics chip, so it reacts as the voice arrives: it looks at you while you speak, looks away while ChatGPT thinks, and moves its mouth as the answer is spoken.
Is the lip sync exact?
It is close and it is live. Facelet listens to the voice it is playing back and animates the mouth and face from it on your Mac. The mouth can trail the voice by a fraction of a second. It is not a pre-rendered video, so it is never out of step by more than that.
Can I control how the character moves?
Yes. Ask for a nod, a shake of the head or a wink from the Gestures menu, or tell it to hold still. Everything else, the gaze, blinks and breathing, happens on its own.