How it works
Last updated 2026-09-03
Upload, set the scene, and start talking. Here is what happens at each step and what it costs.
1 · Bring a face, or don't
A photo or short video works. So does an anime character, a pet, or a description with no image at all — the model generates from the prompt in that case.
If the material shows a real person, you confirm you are that person or have their permission. Until that is on file, the character cannot generate anything.
The photo is what the character looks like: the video model uses it as the reference for every scene.
2 · Say what you want to happen
One sentence is enough. 'A rainy evening in Kyoto, she is an old friend I have not seen in years.' The system turns that into a structured scene — relationship, place, time, goal, tone — which you can edit.
The twelve cards on the homepage are starting points, not a menu. Anything you can describe works.
3 · Say something, and see
A few phrases sit under what she says. They are there to help you find words when you cannot — not branches to choose between. Type whatever you like instead; it goes to the same place.
Nothing is generated ahead of you. What you say enters a world that is already moving, and the world decides what comes of it — which is why there is a delay, and why she does not always hear you. There is one timeline, it only goes forward, and it cannot be re-rolled.
4 · Video, only where it matters
Everyday beats run in text and voice over the last frame the world rendered. That is cheap, so text and voice are generous.
Full video renders the moments worth rendering: turning points, travel scenes, and anything you explicitly ask for. That is the expensive part, so that is what a plan meters.
Nothing is ever rendered for a moment that has not happened. There is no future to render — the world has not computed it yet. That rule is in the code, not the config.
5 · It remembers, and you control that
Facts worth keeping get written to memory and pulled back when relevant — not your whole history, only what matters to this character and this moment.
You do not get a panel for editing what she remembers. That is on purpose — her memory is hers, and a world you can correct from the outside stops being one that runs without you. What you can do is delete the world.
6 · What "world model" means here
Three layers, and they are separate on purpose. The premise is frozen the moment your first sentence is compiled — the place, the language spoken there, the timezone, the season, who she is to you — and every later step is written against it, so the world cannot quietly drift into a different city.
The world's own state is what changes: time, weather, light, the objects in the room and where they are, the doors and floors, the people passing outside. Move a cup and it is somewhere else from then on. Close a window and it stays closed, and the rain gets quieter.
Her state is hers: what she wants, how she feels, what she has come to believe about you. It is not a script and there is no panel for editing it.
The camera is your position in that space — where you are standing, what you are looking at, how far away. It is a first-person view of a place with a floor and a direction, not a video generated to look like one.
None of this is regenerated when you come back. The world keeps its own clock while you are gone and catches up on the gap, so returning after three hours means arriving three hours later — not resuming a paused scene.
Questions about this page? legal@youhere.live