Chinese tech giant ByteDance is developing artificial intelligence (AI) that creates a three-dimensional representation of the real world, responding in real-time to user movements and voice. The company, known for growing TikTok into a global video platform, is entering the so-called 'World Model' competition, leveraging its video generation technology for applications in robotics and autonomous vehicles.
Bloomberg reported on September 7 that ByteDance is preparing a new AI model specialized in real-time spatial video generation, with a potential launch as early as next month. Zhang Yiming, the founder of ByteDance, is personally overseeing the development, which involves personnel and AI resources from various business sectors.
The new AI model will be based on ByteDance's existing video generation AI, Sidance. It aims to create a virtual world where users can interact in real-time, applicable for live broadcasts, short dramas, and games.
This technology, referred to as a World Model, allows AI to learn how the real world moves and changes, going beyond simple text or image generation. It can understand and predict outcomes based on user interactions, such as how moving an object affects its environment or what happens in a space where a car is moving. This capability is considered crucial in fields that require interaction with the real world, such as robotics, gaming, and autonomous vehicles.
ByteDance's strategy is to leverage its technological expertise in video to enter the World Model market. The Sidance video generation AI has already established itself as a significant success for the company, being utilized in various products, including the video editing program CapCut and the AI chatbot Doubao. Independent filmmakers, content creators, and AI startups also use Sidance.
Moreover, the new AI model is expected to connect various businesses within ByteDance. It will integrate AI models and cloud computing resources with content platforms like TikTok and hardware from XR (extended reality) device manufacturer Pico.
For instance, the new AI model could create a virtual world that responds in real-time to the voice or movements of Pico headset users, generating video at 20 frames per second with a delay of about 0.05 seconds.
ByteDance's development of the World Model is seen as part of its strategy to transition its business towards generative AI. The company aims to expand its influence in the AI sector by connecting AI models with data centers, cloud services, content platforms, and XR devices.
To enhance its AI competitiveness, ByteDance is also making significant investments. According to Bloomberg, the company recently secured a $30 billion loan to expand its AI capabilities and acquire data center and AI hardware. ByteDance is reportedly considering increasing its capital expenditures for AI infrastructure to as much as $70 billion this year.
Meanwhile, Meta and Apple have also made substantial investments in virtual and mixed-reality devices. Meta is building a popular ecosystem around its Quest headsets, while Apple is promoting its high-end spatial computing device, the Vision Pro. However, both companies are still facing challenges in establishing VR and spatial computing devices in the mainstream market.
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.

