At the end of 2025, during a dinner with the model team, Zeng Yan, head of the Seedance model, told her manager that she still wanted to try training a larger model, at least one with 200 billion parameters.
Zeng is a young researcher on ByteDance’s Seed team. She joined the company as a campus hire in 2021, and her research has focused on video understanding and generation. “She has always had strong technical judgment,” someone who has worked with Zeng told 36Kr. “She is also proactive. When she believes in something, she is quite persistent and will find ways to fight for the resources to make it happen.”
That assessment closely captures Zeng’s situation during the training of Seedance 2.0. “She wanted to train a model with more parameters, but there were disagreements within the team at the time. Some people thought scaling the model directly to the 200–300-billion level was a bit aggressive. Training resources were also tight, so they thought it might be safer to train a model at around the 100-billion level first,” a person familiar with the matter told 36Kr.
The person said a senior executive happened to attend the dinner. After learning of Zeng’s idea, the executive said resources could be coordinated for the training of Seedance 2.0. “Later, at Seed’s final formal internal review meeting, Seed head Wu Yonghui and Zhou Chang, who leads visual multimodal generation, discussed it and ultimately chose to support Zeng’s idea.”
The technical choice, which seemed aggressive at the time, was later validated. “Seedance 2.0 only succeeded because Zeng insisted that the model had to be large enough and that the training data had to be sufficiently rich,” the person said.
36Kr interviewed multiple artificial intelligence practitioners inside and outside ByteDance. Most said ByteDance’s reputational turnaround in foundation models began with Seedance 2.0. In large language models (LLMs), Doubao was not widely seen as entering China’s top tier until its foundation model iterated to version 1.6 in mid-2025. Its coding capabilities had not made a major breakthrough before Doubao 2.1 was launched, and in China, it still lagged behind the flagship models of startups such as Z.ai, Moonshot AI, and DeepSeek. Seedance 2.0, by contrast, can be considered ByteDance’s first model, and currently its only model, to achieve a clear performance lead.
This leading model was also the first to generate meaningful revenue for ByteDance. 36Kr previously reported exclusively that, since 2026, Volcano Engine has continuously raised its revenue targets for its model-as-a-service (MaaS) business. Tan Daize, head of Volcano Engine, told 36Kr that more than half of Volcano Engine’s MaaS revenue this year has come from Seedance.
ByteDance had found a promising business.
There are different accounts of Seedance 2.0’s profit margin. LatePost previously reported that the model’s gross margin reached 70%, while several AI practitioners told 36Kr that they estimated Seedance 2.0’s gross margin at 90%. Tan, however, said the outside world has paid too much attention to Seedance’s profits, and that the figure is not actually that high. In any case, compared with LLMs at the current stage, video models are a better business in China: they can command higher prices, and competition is less intense.
ByteDance has also built a distinctive cycle across several business lines. Content platforms such as Hongguo and Douyin are giving AI-generated dramas more support than before. More content production companies are buying models from Volcano Engine to produce high-quality AI-generated short dramas. The surge in AI-generated dramas is bringing more advertising or paid traffic revenue to Hongguo and Douyin. A loop is forming.
When model capabilities improve in a step change, they often unlock new productivity scenarios, and the flow of money changes as a result. This was validated once after Claude Opus 4.6 was released. With Seedance 2.0, it has been validated again.
Seedance is not only ByteDance’s first major turnaround after more than three years in foundation models. It has also shown China’s broader foundation model industry something important: Foundation models require enormous capital expenditure, but with sufficiently high margins, they can also make money.
Talent and data drive the turnaround
ByteDance has long been willing to invest heavily in AI. GPUs, data, and talent have all received significant funding.
Talent came first.
ByteDance has long been adept at delegating new businesses to experienced hires, and it initially applied this approach to Seed. But after more than a year, progress was limited. In mid-2024, ByteDance began placing heavier bets on AI talent recruitment. This was the first major change after the core management team became deeply involved in the AI business: hunting foundation model talent around the world, even at extremely high salary levels.
In mid-2024, ByteDance established a recruitment team dedicated to serving the AI business. What made the team special was that many of its human resources staff had strong business backgrounds, enabling them to communicate better with high-level AI talent. Both of its leaders had worked on important businesses at ByteDance and had strategic experience. ByteDance also offered the team high compensation. HR staff working on senior recruiting often earned more than RMB 1 million (USD 147,534) a year.
A talent list also began to take shape. It covered top AI talent in China and overseas. “There are a few hundred people on it, and it is updated continuously,” said an HR staffer involved in senior recruiting. “The Seed department basically does not need to follow ByteDance’s salary system when making offers. As long as we want to recruit someone, the compensation package can be customized.”
Zhang Yiming, founder of ByteDance, also returned to a more proactive role, frequently meeting foundation model talent for “direct chats.” Many high-level researchers chose to join when faced with both money and sincerity.
People are important because, as multiple industry insiders put it, “If an experienced leader who can make the right directional judgments takes a team of smart enough young people to steadily run training experiments, and is given enough resources, it is hard for the result to be bad.”
As Zhou Chang and Wu Yonghui joined in succession, Seed reduced internal competition within the same technical directions, and its judgments on technical direction and training approaches became more convergent and more accurate. This was especially evident in Seedance.
Previously, ByteDance had two independent teams training video generation models. Zeng Yan, at its AI lab, was working on PixelDance, while Jiang Lu, at Seed, was working on Seaweed. “At that time, different camps had formed inside ByteDance. The various video teams initially made the wrong choice in architecture selection,” said a person familiar with the early situation.
Take the first version of PixelDance, for example. It chose an architecture that extended 2D U-Net to 3D, which can be understood as turning an image diffusion model into a video model rather than using a native video architecture. At a time when the technical path for video generation was far from settled, this looked faster and safer. In hindsight, it proved to have a lower ceiling.
During the same period, Kuaishou’s Kling AI team chose the same diffusion transformer (DiT) architecture as OpenAI’s Sora, and it used a native video approach, which had a higher ceiling and was also more difficult. “If people and data do not hold it back, the model’s trained performance will not be poor,” a person familiar with the matter said. The time gap that formed during this period allowed Kling to lead for nearly the following year.
At the end of 2024, Jiang left ByteDance to join Apple. Zeng joined Seed after organizational changes at the AI lab and became the main person in charge of Seedance, the visual generation model. It was a lean team, with only a dozen core algorithm engineers. From that point on, ByteDance’s video generation model team and technical path began to converge.
During the transition from the later stage of PixelDance to Seedance, the biggest change was an adjustment in the training architecture. The team shifted from a U-Net architecture to a DiT-based architectural direction. This architecture is better able to realize the scaling law of foundation models: As parameters, data volume, and computing power grow, model performance improves significantly.
But neither Seedance 1.0 nor 1.5 could be considered successful. Both versions had a clear gap with Google’s Veo. Their global market share ranked around fourth, and in China, they were also behind Kling. A Volcano Engine employee said the rivals ByteDance was watching closely during that period were Kling and Runway. “It fought Kling AI for quite a long time,” the person said, but it still failed to catch up. A marketing head at a video agent company told 36Kr that Kling once held nearly 80% of the video generation market.
The long-running rivalry between ByteDance and Kuaishou had moved from short video, live streaming, and e-commerce to video models.
In December 2025, not long after Seedance 1.5 went online, the team began investing in training version 2.0. That would lead to Seedance’s turnaround.
“Seedance places great importance on training data and model architecture,” a Seed employee said. In a simple comparison of GPU quantities, ByteDance trails OpenAI and Google in both quantity and quality. Overseas companies have generally started using large clusters of Nvidia B-series GPUs, meaning large AI computing clusters built with Nvidia B200 and B300 GPUs, to train video models. Inside ByteDance, although Seedance also has high priority, it still ranks behind language models and uses less advanced GPUs for training. As a result, the team could only focus on architecture and data.
On people and organization, a foundation model entrepreneur who has had contact with the Seedance team told 36Kr:
“Major technological innovation often requires intelligent researchers and a relaxed research environment to stimulate them. But once the technical path is determined, what is needed is strict schedule management to ensure every step of training is executed properly. The Seedance team really did every detail well. Its combat effectiveness as a coordinated force is very strong.”
The element widely regarded as the most critical to Seedance 2.0’s success was the data used for training.
The pretraining stage of foundation models requires many small-scale experiments. It needs many people to handle data standards, synthesis, sourcing, annotation, cleaning, evaluation, and other work. It also requires a sufficient amount of rich and diverse data.
“Seedance’s data annotation was done very robustly. That allows user prompts to be matched very precisely to specific data,” a person familiar with the matter told 36Kr. Inside ByteDance, the team doing model data evaluation for Seedance alone has more than 1,000 people.
By comparison, for many leading startups in the video field, an internal evaluation team of several dozen people would already count as a relatively large investment. Behind each Seedance algorithm engineer, there are often more than a dozen data staff providing support. This enormous data team needs to quickly deliver the data required by algorithm engineers for experiments and training, and collect user feedback to ensure algorithm personnel can quickly capture any signal useful for model iteration.
Within the Seedance algorithm team, there is also a person specifically responsible for interfacing with the data team. Multiple people close to Seed told 36Kr that this role is like a data product manager within a training team and is highly important. The person clearly understands what kind of data the training requires, can make clear and effective demands of the data team, and the algorithm team even constructs data itself and performs refined data cleaning on its own.
In terms of specific data selection, Seedance used almost no Douyin data. Instead, it purchased a large amount of film and television resources, especially large quantities of high-quality cinematic materials. It then used language models to break them down into scripts and storyboards. “The effect is like putting a group of film masters into the model.”
The team conducts targeted data training for different scenarios. For example, “Seedance 2.0 systematically studied film and television works rich in running and jumping movements because they can reproduce many real motion scenarios. The team also simulated or screen-recorded games to learn from them, and used professional scenario data to help the model learn various indoor spaces,” a ByteDance data employee told 36Kr. Data from these various scenarios formed Seedance 2.0’s enormous training dataset.
ByteDance has invested heavily not only in training data for video models, but also in data for its other models. 36Kr learned that, in training world models and coding models, the data budget at the beginning of 2026 exceeded USD 10 million, and “if it seems insufficient, the budget can be increased at any time.”
There is an important reason for this enormous investment in AI data: Seed established a principle from the beginning that it would “not do distillation.” If all data has to be synthesized, purchased, and cleaned by the company itself, a large team is inevitably required for support. “ByteDance’s goal for almost all of its models is to reach the global top tier, or even state-of-the-art standards. That is hard to achieve through distillation,” a ByteDance employee told 36Kr.
Volcano Engine finds a selling point
“The internal momentum is surging,” a Volcano Engine salesperson told 36Kr. Because of Seedance 2.0’s popularity, the Volcano Engine sales team went out in full force. “Everyone is selling Seedance to clients.” When clients hesitated, some salespeople would confidently reassure them of the model’s quality.
In April 2026, Volcano Engine opened sales of the Seedance 2.0 API to all clients on one condition: A client had to sign a Seedance annual usage framework contract worth at least RMB 10 million (USD 1.5 million) upfront to qualify for the full capabilities of 2.0. This refers to the version that can handle high concurrency and supports licensed real-person likeness rights. These two points are crucial for Seedance 2.0’s main user group: AI content production companies that own real-person intellectual property and need to mass produce content.
This approach had not appeared in Volcano Engine’s previous MaaS sales plans. After DeepSeek R1 appeared in February 2025, LLM prices were suddenly driven down, and competition among model providers rapidly shifted into a buyer’s market. For buyers, the choice was simple: Use whichever model is good and cheap. Competing for market share at rock-bottom prices left model providers with almost no pricing power. For Volcano Engine, which mainly sells its own models, it was even harder to set such a seemingly high sales threshold.
“In 2025, Volcano Engine had limited models it could sell and followed a comprehensive cost-performance route,” a Volcano Engine employee told 36Kr. “Competition among language models was too intense, and at the time, a high-value scenario like coding had not broken out. The Seedream image model was fairly competitive in China, but the volume was not large.” Because Seedance 1.0 and 1.5 were not good enough, they could not beat Kling AI in market competition, and few clients paid for them.
“By the beginning of 2026, we had placed all our hopes on Seedance 2.0. We had to come up with a new story to tell clients,” the Volcano Engine employee said. At the time, they had heard internally that Seedance 2.0 was a strong model and thought it might have a lead of one or two months. “But we did not expect it to be this strong. Several months have passed, and our advantage is still there.”
The first wave of feedback that the model was useful came from consumer users. During the 2026 Lunar New Year holiday, ByteDance’s consumer AI products, including Doubao, CapCut, and Dreamina, began fully integrating Seedance 2.0. Traffic poured in like a flood, causing users to wait in line for up to ten hours to generate a 12- to 15-second video.
Dreamina responded quickly by offering a paid option: users who subscribed and bought credits could receive priority in the queue. Users who could not tolerate the wait contributed Seedance 2.0’s first batch of revenue. According to an estimate by a strategy employee at a major company cited by 36Kr, Dreamina’s revenue in March this year was about RMB 140 million (USD 20.7 million), and in April 2026, it reached RMB 210–220 million (USD 31.0–32.5 million). The vast majority came from calls to Seedance 2.0.
But Volcano Engine soon took over the baton for monetization.
One afternoon in March, Zhou Zhipeng was hiking in the mountains with his family, but his phone would not stop ringing. “I took six calls from Volcano Engine salespeople, all trying to sell me Seedance 2.0,” he said. Zhou is a co-founder of content production company Small Design. Two short dramas produced by the company with the help of AI, The Golden Tomb Seeker and The Hunger Tower, were showcased at the Cannes Film Festival in May.
Because his company has an AI creation tool platform, Zhou had been dealing with Volcano Engine salespeople for more than a year, but he had never placed an order. For leading content production companies, spending RMB 10 million a year on AI models and tools is not a large sum. What made him hesitate was that spending that RMB 10 million would effectively bind his company to ByteDance’s video generation model for one year. If another model better than Seedance 2.0 and its later versions emerged within that year, it would be difficult for him to migrate quickly because the money would already have been spent on Seedance.
That meant betting on how far ahead ByteDance was in video generation models, and on how long its lead could last. Many short drama practitioners told 36Kr that Seedance 2.0 is better than existing video models in multi-shot narrative, consistency across shots, and synchronized audio-video generation. These capabilities are highly important for industrialized content production.
“The day after I got back from the hike, I decided to sign with Volcano Engine,” Zhou said. In his view, competition in AI-generated dramas accelerated after the Lunar New Year holiday. Whether a company could use the best model at the earliest possible moment determined whether it could be among the first to launch high-quality AI-generated dramas in the market. That affected competitive speed, and even survival. “Once we decide to put production capacity and tools on the Seedance model, annual consumption will definitely be more than RMB 10 million. It could be RMB 20–30 million (USD 3.0–4.4 million).”
Zhou was not alone in thinking this way. Leading short drama and comic-style drama production companies such as COL Group, Jiuzhou Wenhua, and Jiangyou Wenhua all paid up. Some larger payments reached RMB 50 million (USD 7.4 million) in a single top-up. A group of smaller companies also tried to pool orders to reach RMB 10 million so they could squeeze onto this fast-moving train.
From a pricing perspective, compared with language and speech models, the unit price of video models can be dozens of times higher. According to official pricing, generating 720p video with Seedance 2.0 costs about RMB 1 (USD 0.15) per second. That figure is nearly double the price of domestic video models of the same generation. Kling AI quickly turned price into one of its advantages, using discounts of 20–30% to win clients, while models such as Shengshu Technology’s Vidu attracted comic-style drama and anime-style content clients that cared more about cost performance.
Although Seedance 2.0 is currently the most expensive model in the industry, it does not offer discounts, even to affiliated businesses such as CapCut, Volcano Engine, and Hongguo. Externally, Volcano Engine salespeople also bundle products depending on each client’s willingness to buy.
But video production companies with high requirements for video quality and scaled production have chosen to accept it. “Last year, Volcano Engine salespeople were begging us to buy. This year, we are begging them to sell,” another practitioner said. “The results of 1.0 and 1.5 were not good. Over the course of a year, we might have called the model only a few hundred times, and most of our spending was on Kling. But after using 2.0 this year, we found that it really has no rival. We have to go all in.”
Content production companies have poured in heavy investment. It again shows that a globally competitive model can capture money from the broader market.
The growth of Volcano Engine’s MaaS business has put enormous pressure on salespeople at rival companies. “Volcano Engine is the toughest opponent Alibaba Cloud has ever faced,” an Alibaba Cloud salesperson told 36Kr.
In the previous era of cloud computing, Alibaba Cloud had been far ahead for nearly a decade. But the arrival of the foundation model era has created a new incremental market outside public cloud, centered on GPU cloud computing power. With its breakthrough in video models, Volcano Engine has found a new growth curve.
To lift revenue faster, Alibaba Cloud separately formed several MaaS sales teams starting at the end of 2025, specifically to pry open existing large cloud computing clients. It also sharply increased the weight of token sales in sales key performance indicators. Call volume multiplied by an incentive coefficient was included in performance, and salespeople were urged to help clients think through “what scenarios could actually use tokens.”
But for Alibaba Cloud, client scenarios are not the most crucial factor. “For MaaS to sell well, the key is still that model capabilities must be strong enough,” the Alibaba Cloud salesperson said.
Volcano Engine, meanwhile, has recently begun to focus on another metric: market share. This year, it wants models across different modalities to increase their market share in regions around the world. Seedance 2.0 currently ranks second globally in market share, behind only Google’s Veo, which holds nearly half of the global market. In the second half of the year, with Seedance 2.5 set to be released soon, Volcano Engine’s goal for Seedance is clear: take the top spot globally.
Hongguo and Douyin complete the loop
The revenue Seedance brings to Volcano Engine is direct and explicit, but ByteDance’s other business segments are also benefiting from it. The video model has formed an almost ByteDance-exclusive commercial closed loop with the company’s various business segments.
Hongguo Short Drama and Douyin are the most important nodes in this loop. They absorb the large number of AI-generated short dramas made with Seedance and convert them into different forms of revenue.
Since the beginning of 2026, Hongguo has added AI-generated dramas to the list of content it is focused on supporting, alongside premium dramas.
The most direct difference is reflected in revenue sharing. “The revenue-sharing multiplier for comic-style dramas is 40–50 times, higher than in 2025, while photorealistic AI dramas can reach 60–80 times. The revenue-sharing multiplier for nonpremium live-action dramas has dropped to 40 times, although the revenue-sharing base for live-action dramas is still higher than that for AI-generated dramas,” the head of an AI drama production company told 36Kr.
Behind the difference in revenue sharing is an enormous difference in production costs moving in the opposite direction: The formats with higher revenue sharing actually have lower costs. Wang Xiaoshu, founder of short drama distributor Jiashu Technology, roughly calculated the numbers: A short drama of around 100 minutes, excluding top-tier premium dramas, costs RMB 500,000 (USD 73,767) to RMB 1 million to shoot in live action, but only RMB 50,000–100,000 (USD 7,376.7–14,753.4) to produce with AI.
Stimulated by both lower costs and stronger incentives, short drama production companies have stopped most of their live-action projects and shifted to AI-generated dramas. This change has not only brought Volcano Engine a large number of clients coming specifically for Seedance, but has also directly brought Hongguo a large supply of content, ultimately leading to higher daily active users, user time spent, and revenue.
All short dramas that cooperate with Hongguo are launched simultaneously on Hongguo and Douyin. One industry insider roughly calculated that last year, the entire industry received about RMB 1 billion (USD 147.5 million) in revenue sharing each month from live-action dramas. This year, only two or three months after AI-generated dramas began to grow rapidly, monthly revenue sharing has reached at least RMB 300 million (USD 44.3 million).
Beyond advertising revenue sharing, there is another, larger pool of money.
36Kr learned that Hongguo does not allow distributors to buy traffic placement. The platform relies entirely on algorithms for distribution, and its support for top premium dramas, including AI-generated dramas, is mainly at the marketing and promotion level. As a result, short video platforms, including Douyin, Kuaishou, and WeChat Channels, have absorbed large amounts of paid traffic spending.
“Paid traffic is the most important way for AI-generated short dramas to acquire users across content platforms,” Wang said. And because content production costs have fallen sharply, budgets that can be used for traffic acquisition have increased significantly.
“At the beginning of the year, our daily spending on AI-generated dramas may have been only a few hundred thousand RMB. Now it has reached several million RMB a day,” Wang told 36Kr. And that is only the spending scale of his company. “The industry’s overall paid traffic spending, including for live-action dramas, is in the billions of RMB each month and has grown by a two-figure percentage. The main incremental growth is in AI-generated dramas.” These expenses are being split among video platforms such as Douyin and Kuaishou.
The same applies to marketing companies that place in-feed ads on Douyin and to the advertisers behind them. Because using Seedance to make advertising materials is easier than live shooting or animation, costs fall and efficiency rises. Without changing the budget, advertisers can produce more ad materials and allocate more money to ad placement. “Ocean Engine will ask us to help major clients on its platform use AI to mass produce ad materials. The goal is also to help big clients increase ad spending,” a marketing practitioner using Seedance 2.0 told 36Kr.
Video models can also make money
Foundation models are expensive to build.
Training costs alone can range from nine- to ten-figure sums in USD. As model sizes continue to grow, research by Epoch AI shows that, over the past several years, the training cost of each generation of LLMs has risen by two to three times. Early this year, Anthropic CEO Dario Amodei said in an interview with Time magazine that the training cost of Anthropic’s next-generation AI model would reach about USD 1 billion, and that the following generation could climb to USD 10 billion.
That does not include the inference costs incurred every time users use a model after it goes online. “The larger the user base, the larger the losses” was once the dilemma facing chatbot products. As early as the initial release of ChatGPT, SemiAnalysis estimated that once a model reached a certain user scale, weekly inference costs would exceed the total cost of a single training run.
That judgment has been repeatedly validated over the following years. According to Sacra’s estimates, OpenAI’s inference costs reached USD 8.4 billion in 2025 and are expected to climb to USD 14.1 billion in 2026. That is despite the cost per token falling by more than 99% within two years, while OpenAI’s revenue over the same period was only more than USD 20 billion.
As questions grow over whether foundation models, one of the most capital-intensive businesses in technology, can really make money, Anthropic became the first to dispel some of those doubts. At the end of 2025, Anthropic’s full-year revenue was only USD 9 billion. But just four months later, driven by the Claude Opus series and the explosive growth of Claude Code among developers, Anthropic’s annual recurring revenue had surged to USD 47 billion by May this year.
This change illustrates that models can make money, but they need to meet two conditions: The model must provide high value to users, such as in coding and agents, meaning tokens need to be more valuable. At the same time, the model must be good enough to reach state-of-the-art performance in its field.
The Silicon Valley experience also applies in China. The difference is that, in China, the first breakthrough came from a video model.
There is a certain inevitability to this. China started relatively late in LLMs, but in video models, it began exploring at almost the same time as Silicon Valley. That means everyone stood at the same technical starting line. In addition, China’s content creation space has always been active, and the prevalence of short videos has created a large creator ecosystem. This makes companies such as ByteDance and Kuaishou, which started in short video, naturally more sensitive to opportunities in video models. Video practitioners also better understand how to turn models into products and more quickly close the loop on commercialization.
Tan Daize, head of Volcano Engine, once analyzed the pricing mechanics of video models for 36Kr:
“Your pricing needs to ensure that the client’s migration benefits are at least two or three times higher than the migration costs before people are willing to use it. For example, if shooting an advertisement used to cost RMB 100 (USD 14.8) per second, and now Seedance can produce a similar effect for only a few RMB, then the model has created enough value.”
High value brings greater pricing power, but inference costs for video models can still be further reduced. An entrepreneur told 36Kr that video generation is a compute-intensive task. Unlike language models, it does not have such high requirements for memory bandwidth, so cheaper Chinese-made chips can be used for inference. Not being constrained by the supply of high-end GPUs from Nvidia and others also allows model companies to expand business scale more freely.
These conditions naturally lead to higher gross margins and faster amortization of the costs invested during the training stage.
That can make video models an attractive business. According to multiple practitioners, Kling AI has long maintained a relatively high level of profitability, while MiniMax’s Hailuo video model has also long been the main driver of the company’s overall revenue and maintains positive gross margins.
But after Seedance 2.0 was released, the landscape changed significantly. A leading video agent company told 36Kr that, before the release, video content producers still spread usage across multiple models, with Kling holding the largest share. But after Seedance 2.0 was released, the situation quickly reversed. On the platform, calls to Seedance 2.0 approached 70%. “We have now put 90% of our production on Seedance 2.0,” another leading short drama company told 36Kr.
36Kr learned that, during 2025, Kuaishou’s senior leadership had discussed the goals for Kling many times. At the time, the conclusion was to treat commercialization as an important goal, which inevitably led to lower investment in the next-generation model. Seedance’s sudden rise made Kuaishou realize that “if you want to build the strongest model, the upfront investment must be huge. Making profitability the core goal will very likely slow model development,” a former Kuaishou insider said. This is also the main driver behind Kling’s current independent spin-off and push for an IPO: seeking more space and resources.
The reality is harsh: to turn foundation models into a sustainable business, a company must sustain its model’s lead. Once a model’s capabilities are overtaken, market share will quickly be siphoned away.
36Kr learned that Seedance’s next-generation model, 2.5, will be officially released in late July, while new versions of Kling and Veo will also arrive soon. In addition to large companies, many Chinese video model startups, such as AIsphere and Sand.ai, are also releasing new models at a steady pace.
If ByteDance wants to keep the commercialization flywheel of foundation models spinning at high speed, it still needs to fill in the missing piece: coding capabilities. That has placed Doubao 2.1, which went online at the end of June, in a crucial position. Volcano Engine’s Tan said it has already earned a seat at the table in coding and agents. “There are not many players in China that have truly earned a seat at the table yet,” he said. “Although Seedance sells a bit more now, I hope LLMs can become the main force later.”
KrASIA features translated and adapted content that was originally published by 36Kr. This article was written by Zhang Yuxin and Deng Yongyi for 36Kr.
Note: RMB figures are converted to USD at rates of RMB 6.78 = USD 1 based on estimates as of July 21, 2026, unless otherwise stated. USD conversions are presented for ease of reference and may not fully match prevailing exchange rates.

