StoryResearch
Google DeepMind publishes Genie, a model that turns images into playable 2D worlds, learned from game videos
Google DeepMindGenie world models
Google DeepMind described Genie, an 11-billion-parameter model trained without labels on internet videos that can turn a text description, photo or sketch into a controllable virtual world. It worked out the possible actions by itself, so users can steer the generated scenes frame by frame although the model was never told which buttons were pressed.
- 11B foundation world model trained in unsupervised manner using unlabeled internet videos.
- Comprises spatiotemporal tokenizer, autoregressive dynamics model, and latent action model.
- Generates action-controllable virtual worlds from text, images, photos and sketches.
- Learned latent action space enables agents to imitate behaviors from unseen videos.