Sora Research Preview
Model Overview
Sora is an AI model that can create realistic and imaginative scenes from text instructions. The video model is capable of generating entire videos all at once or extending generated videos to make them longer, up to a full minute of high-definition video.
Sora uses a diffusion model architecture, generating a video by starting off with one that looks like static noise and gradually transforming it by removing the noise over many steps. Similar to GPT models, Sora uses a transformer architecture, unlocking superior scaling performance.
Capabilities
- Long-form Video Generation: Can generate up to 60 seconds of video while maintaining high visual quality and adherence to the user's prompt.
- Complex Scene Construction: Sora is able to generate complex scenes with multiple characters, specific types of motion, and accurate details of the subject and background.
- Physical Understanding: The model understands not only what the user has asked for in the prompt, but also how those things exist in the physical world.
- Image-to-Video: Sora can also take an existing still image and generate a video from it, animating the image's contents with accuracy and attention to small detail.
Official Examples
These examples are generated directly by Sora, demonstrating its capability to follow complex prompts accurately:
Tokyo Walk
Prompt: "A stylish woman walks down a Tokyo street filled with warm glowing neon and animated city signage. She wears a black leather jacket, a long red dress, and black boots, and carries a black purse. She wears sunglasses and red lipstick. She walks confidently and casually. The street is damp and reflective, creating a mirror effect of the colorful lights. Many pedestrians walk about." (Source: OpenAI Sora Announcement)
Paper Airplanes
Prompt: "A flock of paper airplanes flutters through a dense jungle, weaving around trees as if they were migrating birds." (Source: OpenAI Sora Announcement)
Wooly Mammoths
Prompt: "Several giant wooly mammoths approach treading through a snowy meadow, their long wooly fur blows lightly in the wind as they walk, snow covered trees and dramatic snow capped mountains in the distance, mid afternoon light with wispy clouds and a sun high in the distance creates a warm glow, the low camera view is stunning capturing the large furry mammal with beautiful photography, depth of field." (Source: OpenAI Sora Announcement)
Historical California
Prompt: "Historical footage of California during the gold rush." (Source: OpenAI Sora Announcement)
Intended Use & Limitations
OpenAI has noted several current limitations with the Sora model:
- Physical Accuracy: It may struggle with accurately simulating the physics of a complex scene, and may not understand specific instances of cause and effect (e.g., a person might take a bite out of a cookie, but afterward, the cookie may not have a bite mark).
- Spatial Details: The model may confuse spatial details of a prompt, for example, mixing up left and right.
- Trajectory Over Time: It may struggle with precise descriptions of events that take place over time, like following a specific camera trajectory.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity.
Key Features
Generates videos up to 60 seconds in duration
Maintains character and style consistency over time
Simulates basic physical interactions and object permanence
Handles complex camera motions natively
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
Undisclosed
Context Window
undisclosed
License
Proprietary
Deployment
Resources & Links
Lineage
Model Family
Part of the Sora family
Only release in this line currently tracked.
Curator Notes
Sunset: Web/App discontinued April 26, 2026; API scheduled for shutdown September 24, 2026 per official OpenAI Help article 20001152.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model