AIコンサル

Is Losing to Frontier Models Inevitable? Kimi K3's Comeback, the Limits of Scaling, and What Layer Strategy Tells Us About the AI Race Ahead

Published2026-07-26Ryuta Hamamoto

Is it really true that "the leading frontier models take everything and latecomers have no choice but to lose"? This piece breaks down, in plain language, why Moonshot AI (Kimi), a supposed latecomer, clawed its way up to world-class level, what the limits of scaling laws (the data wall) mean, what the next stage looks like (synthetic data, reinforcement learning, test-time compute, autonomous agents), and finally the "which layer do you fight in" question of layer strategy. Kimi K3 sits in the top tier for cost-effectiveness on the benchmarks (as reported), and in my own firsthand use I rate it as world-class. I also share, as my own view, where Japanese companies and startups might build a durable position.

Is Losing to Frontier Models Inevitable? Kimi K3's Comeback, the Limits of Scaling, and What Layer Strategy Tells Us About the AI Race Ahead
シェア

Hello, this is Ryuta Hamamoto from TIMEWELL.

When you look out over the world of AI, you often hear a certain anxiety. "In the end, won't the giant models running out front, the so-called frontier models, take everything?" "For companies that arrived late, or startups with little capital, isn't it a contest before it even begins?" I hear this question again and again, both from executives whose companies work with AI and from developers about to build something.

Let me give you my conclusion up front. I do not think losing is inevitable. The reason is that the rules of the competition themselves are being rewritten right now. In this article I will look at why, from three angles in turn. The first is a concrete case: why China's Moonshot AI, the company behind the Kimi models and a supposed latecomer, clawed its way up to world-class level in such a short time. The second is the story of how the "scaling law" that has made AI smarter until now is hitting a wall. And the third is where the axis that decides winners moves once you are past that wall, and the "which layer do you fight in" question of layer strategy.

I will unpack the technical terms as they come up, so I have written this to be readable even if you are not steeped in AI technology. For the parts that are matters of fact I attach sources, and for the parts that are a view about how the future might go, I mark them clearly as "my view" and keep them separate. If you are wondering where your own company's AI adoption stands right now, it helps to first check your footing with our free AI literacy self-check. The strategy discussion in the second half will feel much more like your own concern once you have.

Is "the frontier takes it all and you lose" really true?

First, let me share the idea that sits at the core of this whole article.

The reason many people feel that "the winners of AI are already being decided" is that they hold one assumption. That assumption is that the strength of an AI is decided almost entirely by "how much data and computing power you can pour in." Stand on that assumption, and it does follow that a small number of players with the most capital and compute have the advantage, and that latecomers and the lightly funded have no chance.

The important thing here is that this assumption is not correct forever. The rules of competition get rewritten when the phase of the technology changes. There was once an era when people believed "the company with the biggest factory wins manufacturing." Yet with the outsourcing of design and parts, and the rising importance of software, the way to win changed many times over. The same thing is about to happen with AI. The axis is shifting from an era of pushing through with "bigger, more" to an era where the gap opens up on "how efficiently, how cleverly you can run it."

I will confirm this "the axis is shifting" phenomenon with two concrete examples. One is the fact of why Kimi, supposedly a latecomer, caught up to world-class level. The other is the structural story of why the old winning pattern, "feed it huge amounts of data to make it smart," is nearing a ceiling. Precisely because these two are happening at the same time, I do not think losing is inevitable.

For what it is worth, this theme connects deeply with a piece I wrote earlier, on open weights and American AI leadership. That one dealt with "what it means to open a model." Think of this article as a companion piece that goes a step further and asks "on an open foundation, how does a latecomer win anyway."

How to read Kimi (Moonshot AI)'s comeback

Now let us get into the first concrete example. This is the story of Kimi, a model from China's Moonshot AI, a Beijing-based company rendered in Chinese as Yuezhi Anmian (the dark side of the moon).

What kind of company is Moonshot AI?

Moonshot AI is a relatively young startup, founded in March 2023 by Yang Zhilin and others1. In China, a handful of AI startups with momentum are sometimes grouped together as the "AI Six Little Tigers," and Moonshot is counted among them. Its funding has included a round led by Alibaba, and as of March 2026 it was reported to be considering a listing on the Hong Kong Stock Exchange1. To be clear, what I am introducing here is neither a "Chinese companies are amazing" story nor, conversely, a "we should be wary" story. I am looking at it neutrally, simply as one case of a latecomer player catching the change in the rules of competition well and staging a comeback.

What the latest model, Kimi K3, showed

Moonshot has kept shipping models at a remarkable pace over a short span. Its first open-source LLM, Kimi K2, came in July 2025, followed by Kimi K2.6 in April 2026, the code-specialised K2.7 Code in June 2026, and the latest, Kimi K3, announced on 16 July 20262. In roughly a year, it has stacked generation upon generation. This sheer speed of iteration is itself a major weapon for a latecomer chasing down a front-runner.

The latest Kimi K3 is technically quite ambitious too. It holds a scale of 2.8 trillion parameters, yet rather than running all of them every time, it activates only the subset of relevant specialists needed for each input, an arrangement called "sparse MoE"2. MoE stands for Mixture of Experts. To put it by analogy, it is like having 896 specialists on staff and calling in only the 16 relevant to each case to do the work. That is far more efficient than always running everyone, while still preserving the cleverness that comes from large scale. On top of that, K3 is natively multimodal, handling not just text but images and more from the outset, and its context window, meaning the length of text it can take in at once, reaches one million tokens2. In the sense of handling long documents and long sequences of work all together, this bears on the autonomous-agent use cases discussed later. Moonshot describes it as "the first open model to reach the 2.8 trillion parameter scale"2.

The fact that its cost-effectiveness is top-tier

Here is the heart of the matter. What is impressive about Kimi K3? In a word, it is the standout cost-effectiveness, meaning "cleverness per dollar you pay." This is not my impression; it is a fact backed by third-party benchmarks.

One yardstick for measuring the cleverness of models side by side is the Intelligence Index published by a third-party firm called Artificial Analysis. Kimi K3 records a score of 57 on this index, a level well above the average for models in the same price band3. To put a more concrete number on it, using K3 costs, when averaged out, about 2.31 dollars per million tokens. A token is the small unit AI uses when it processes text; here, just think of it as "the price for making it do the same amount of work." Claude Fable 5, a leading representative of the top tier, scores around 60 on the same Index, edging out K3 slightly, yet its unit price is reported at about 7.70 dollars per million tokens, more than three times K3's3. You get nearly the same cleverness at close to a third of the unit price. Measured by the yardstick of intelligence per dollar paid, K3 sits in the top tier, and that is the picture the third-party benchmark shows.

On the ranking of raw cleverness itself, the position moves over time, so let me write it precisely. At its debut, K3 came in third on the Intelligence Index, behind Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol, and was reported to be on par with Opus 4.8 and GPT-5.53. Afterwards, as more comparison models were added, its rank on the detail page shifted to seventh out of 190 models3. So it is not accurate to write "always the world's smartest." Precisely put, the true picture is "it broke into the world's top three at debut, and it holds a top-tier position on cost-effectiveness." On AA-Briefcase, a benchmark that measures agentic knowledge work, it was reported to place second behind Fable 5, and on a separate benchmark measuring web development it was reported to have surpassed its rivals3.

Here, let me also be honest about my own experience of using it. I use a range of models side by side in daily practice, and my firsthand impression from actually working with Kimi K3 is that it is world-class. Its persistence when I hand it coding or research tasks, holding a long context while responding to intricate instructions, is, frankly, on a par with the best closed models. This is my own personal assessment based on firsthand use, and I keep it distinct from the "top benchmark score." The top of the benchmark is held by the likes of Fable 5, and I separate the two here so as not to confuse them.

For the sake of balance, let me also touch on the weak points. K3 is a reasoning model, meaning a type that uses a lot of tokens to "think" before answering, so its generation speed, at around 32.1 tokens per second, is not fast, and the wait for the first answer to come back is somewhat long too3. It also tends toward verbose output; when Artificial Analysis ran the same evaluation, K3's output token count reached roughly twice the median of other models3. Even if the unit price is low, using many tokens on a single task drives the actual cost up accordingly, so cost-effectiveness has to be assessed on a per-use-case basis, averaged out. For uses that put speed, or lean and concise output, first, these characteristics do not always fit. Rather than praising it wholesale, weighing its strengths and weaknesses and using it accordingly is, I believe, the stance of a practitioner.

Why could a latecomer catch up to world-class level?

So why was Moonshot, supposedly behind in both capital and compute, able to close the gap this far? Let me offer my own analysis of several factors.

First, the open-weights strategy. Open weights refers to publishing the weights (parameters) that are the inner contents of a trained model, putting it in a state where anyone can download it and run it in their own environment. Moonshot published its flagship models and pulled in the developer community and ecosystem in one stroke. The first-generation Kimi K2 was released under a relatively permissive modified MIT license, and the day after release it was reported to have become the most-downloaded model on Hugging Face, a model-sharing site4. For the latest K3 too, Moonshot has announced that it will publish the weights by 27 July 2026 (as of this writing, it is still awaiting release)2. It has ridden well on the larger current of Chinese players countering US technology controls with openness.

Second, efficiency in training and inference. K3 uses a technique called Kimi Delta Attention, or KDA for short, to curb memory use while speeding up generation when handling long contexts. And at the K2 stage, the company stated in its technical report that an optimiser-family method called MuonClip let it train efficiently and stably even with fewer tokens2. This part gets technically deep, so please take it as report-based information. The gist is "they stacked up refinements that draw more out of the same amount of compute."

Third, the cost advantage. As already noted, it delivers comparable-tier cleverness at a fraction of the unit price of the major frontier labs. The cost structure of development and operations within China likely contributes to achieving this price too.

Fourth, a design that emphasises agents and reinforcement learning. It combines large-scale reinforcement learning (a learning method that gets smarter through trial and error based on rewards from the environment) with synthetic data, optimising for long-horizon coding, knowledge work, and deep reasoning. In K3's demo, it edited a teaser video by itself from 56 source clips, showing agent ability that not only answers but "plans its own steps and sees the job through"2. This direction connects straight to the "coming axis of competition" I discuss later.

Fifth, refinements that cut waste. By splitting off a code-specialised model like K2.7 Code, it avoids running parameters unneeded for that use and raises efficiency2. On top of that, the fast iterative development of releasing K2, K2.6, K2.7, and K3 back to back at short intervals, and the differentiation from a long one-million-token context window, are working in its favour too.

To sum up, Moonshot's comeback comes, in my reading, from stepping off the "build the single largest model" contest and bringing the fight onto a different ground: "open it up and spread it, earn on efficiency, and see long jobs through with agents." This is exactly the living example of "the rules can change," the core of this article.

Looking for AI training and consulting?

Learn about WARP training programs and consulting services in our materials.

The limits of scaling laws and the data wall

Let me move to the second angle. Why are the rules of competition changing now? At the root of it lies a structural circumstance called "the limits of scaling laws."

What were scaling laws?

Behind the rapid rise in AI's cleverness over the past several years lay a simple but powerful rule of thumb. Make the model bigger, increase the data used to train it, and pour in more computing power, and performance rises in step. This is called a scaling law. It is, in a sense, a very easy-to-grasp rule of "make it bigger and it gets smarter," and it was the foundation of the assumption that "those with capital and compute win."

But this rule has a premise you cannot overlook. The data to increase is not infinite.

The wall called the data wall

The pre-training of AI, meaning the first large-scale training in which it acquires its foundational cleverness, has until now been supported by the vast amount of web text humans have written. News articles, encyclopedias, forum posts, blogs, all of it was teaching material. Yet the concern that this high-quality, human-generated text will eventually run short is taking on a real edge. This is called the "data wall."

The analysis that examined this wall quantitatively is the work of a research institute called Epoch AI5. Their research estimated the total stock of effectively usable public text at roughly 300 trillion tokens. Giving it a range, that is 100 trillion to 1,000 trillion tokens at a 90% confidence interval. And when training proceeds optimally from a compute-efficiency standpoint, they project that at an 80% confidence interval, this stock could be exhausted at some point between roughly 2026 and 2032 (centred around 2028). If you over-train, running the same data through learning many times, that timing moves even earlier5. What I want you to note here is that this projection is not a single flat assertion that "it will surely run dry in year X," but a range with breadth. In fact, Epoch AI's earlier estimate put the exhaustion point sooner, but as it became clear that carefully filtered web data can be used in larger quantities than previously assumed, and that running the same data through training multiple times works, the timing has been pushed back5. In other words, there is still room for ingenuity, and yet the wall is drawing near for certain. That is the reality.

This point is also a story about how data is becoming an ever scarcer and more valuable resource for AI. On the value of data itself, and the risk of obsolescence that is its flip side, I dig in from a different angle in my article discussing the risk that AI training data becomes commoditised, so I hope you will look at that too.

The observation that "the era of pre-training is ending"

Researchers who represent the industry are voicing the same concern behind this data wall. Ilya Sutskever, known as a co-founder of OpenAI who later went independent, was reported to have said, in a lecture at the NeurIPS conference in 2024 (not 2026), that "pre-training as we know it will undoubtedly end"6. He is further reported to have expressed the structure in words like "we have reached peak data; there is only one internet, so there will be no more of it" and "data is the fossil fuel of AI"6. The fossil-fuel metaphor is apt, with its implication that reserves are finite and will one day be mined out. Note that although the primary source for these remarks is the video of the NeurIPS 2024 lecture, this article confirms them on a reporting basis, so please read them with that in mind.

As for what comes next, Sutskever was reported to have suggested a move toward the more "agentic" and the more "reasoning"6. This connects straight to the theme of the next chapter.

The next stage beyond the wall

When you hit the data wall, what becomes important for making AI smarter still? The paths being discussed are broadly as follows.

The first is synthetic data. If human-written text runs short, the idea is that AI itself, or some other mechanism, generates new data to use for training. The second is reinforcement learning, where, especially for tasks whose answers can be verified as correct, methods that give rewards and train on them, and methods that polish through feedback from humans or AI, are emphasised. The third is what is called test-time compute, or inference-time compute, which is a way of using a lot of tokens to "think it over" not at the training stage but at the very moment of answering. The earlier point that Kimi K3 is a reasoning model is exactly this. The fourth is autonomous agents, a direction in which AI generates new data and experience for itself while actually interacting with an environment. On top of these, multimodal learning, which handles several kinds of information, and improvements in data efficiency, which learn more from less data, are also raised.

What these have in common is that they steer toward "generating new data and experience" rather than "gathering more of the data that already exists." Because this point is decisive, I will devote a fresh chapter to digging into it.

The coming axis of competition is "cheap, autonomous, efficient"

The third angle, and the thing this article most wants to say. From here on, please read this less as an ordering of facts and more as my own analysis.

Why cost-efficiency becomes the main battlefield

With the data wall, AI's room to grow shifts from "gathering data" to "generating data and experience." Let me pause and think about what this change means.

Synthetic data, reinforcement learning, test-time compute, autonomous agents: pushed to the core, they all converge on a single point. That point is the act of the model itself interacting with an environment to produce new data carrying a reward signal. It explores through trial and error, runs tools, verifies results, plays against itself, and simulates environments. Through these activities, the model acquires for itself "experience not written in the textbook."

Here is where cost per unit starts to bite. When you have to run autonomous agents in large numbers, and for long stretches, to obtain new data, the more you run them the more it costs. The cheaper the inference cost per token, per step, the more trials you can run on the same budget. Conversely, when the overhead is large, from wasted tool calls, overly verbose thinking, and redoing after every failure, you get less experience for the same money.

In other words, the axis of competition shifts from "who has the largest intelligence" to "who can do cost-effective autonomous execution." That is my reading. How cheaply, cleverly, and autonomously can you run agents to generate new data? The scarcer data becomes, the more I believe this single point comes to decide winners.

Why Kimi could be structurally advantaged in this context

Seen this way, the traits of Kimi K3 from the previous chapter suddenly take on meaning. A model that can deliver comparable-tier cleverness at a far lower unit price is, quite simply, at an advantage in uses that run large volumes of agent execution. The fact that, on Artificial Analysis's evaluation, it can produce cleverness just three points behind Fable 5 at around 70% lower cost3 means "on the same budget, you can run more trials than the competition." On top of that, a long context window of one million tokens is a differentiator when going after tasks that require long processes, the so-called long-horizon tasks.

Just to be sure I stress it, this is not an assertion that "Kimi will surely beat the others." I have no intention of pronouncing on which company is superior, and OpenAI, Anthropic, and NVIDIA each hold their own powerful assets and strategies. What I want to say here is the general point that "amid a structural change in which the axis of competition shifts to cost-efficiency and autonomous execution, a model that goes all in on cost-effectiveness stands to catch a structural tailwind." And that tailwind can blow for latecomers and lightly funded players too. That is exactly why losing is not inevitable.

How to embed AI agents into your own operations, and which model to run at what cost: this is a domain where management judgment is tested. In WARP, the AI consulting service we run, the design of which model to run, where, and in what cost structure is always the first fork in the road. In particular, for those who want to roll up their sleeves and launch a business themselves from here, we also offer a program called WARP Entre that stays close to practice.

Layer strategy: which layer do you fight in?

So far we have looked at two pillars: "the rules can change," and "the axis of competition shifts to cost-efficiency and autonomous execution." Finally, let me introduce a framework for answering the question of where, concretely, your own company should fight. Layer strategy, meaning the idea of "thinking in layers."

Huang's framing of "full stack"

There is a figure who has repeatedly put out the view of grasping AI in layers: Jensen Huang, the CEO of NVIDIA. He has spoken of AI and accelerated computing not as a standalone chip but as a "full stack" of several layers piled up, and within the framing that "the data center itself is the new unit of compute"7. It is the idea that not a single computer but an entire building behaves as one giant machine.

Let me add one honest caveat here. In this round of research I could not identify a primary source that lets me assert "it is strictly N layers." So please take the following as an organisation of a general AI stack. I am not asserting the names or order of the layers as anyone's official statement. Generalised and laid out, it comes out roughly as follows.

At the very bottom is the chip, the silicon layer. The GPUs and CPUs that are the heart of computation belong here. Above that is the system layer, the part that integrates equipment at the server and rack scale. Higher still is the networking layer, which links many computers at high speed so that they behave as if they were a single giant machine. Next is the software layer, the base for actually drawing out the hardware's performance. For NVIDIA, the software suite called CUDA belongs here, and it is said to be a powerful barrier to entry, a so-called moat. Above that is the model, the AI foundation layer, which holds foundation models and the mechanisms for running them. And at the very top is the application and agent layer. The actual business applications of each industry, and the agents that autonomously carry out work, create value here.

Vertical integration, or horizontal expansion?

Once you have this layer framework, the differences in each company's strategy come into sharp relief.

One way of fighting is vertical integration. It is a way of bundling several layers in house and turning into weapons the performance and efficiency that come when the layers mesh together, and the switching cost that makes it hard to move elsewhere once you start using it. NVIDIA is said to make both performance and stickiness its strengths through vertical integration that bundles everything from chips to networking to the CUDA software in house7. The more development assets you have that are accustomed to CUDA, the harder it is to move elsewhere. That is the structure.

The other way of fighting is horizontal expansion. It is a way of specialising in one layer and spreading that layer sideways across many customers and use cases. The companies that make models specialise in the model layer and compete horizontally. Cloud operators expand horizontally in the system and infrastructure layers. NVIDIA itself, in the sense of opening CUDA and its systems to a wide range of customers, also has a side that spreads a platform horizontally.

Which layer to fight in, and how far to integrate vertically. This is exactly the fork in the road for each company's strategy. Here too I have no intention of judging companies as superior or inferior. Vertical integration has the strength of tight meshing, but also the heaviness of carrying it all. Horizontal expansion has agility and versatility, but also the risk of being squeezed between others and having your value stripped away. The point is that choosing which layer, and which way of fighting, in light of your own resources and strengths is what matters. That is my view.

And the earlier point, that "the axis of competition shifts to cost-efficiency and autonomous execution," bites on this layer picture too. Whichever layer you stand on, the question within that layer becomes "can you produce the same result more cheaply, more efficiently?" The choice of layer, and efficiency within the layer. Designing these two at once will, I believe, be the backbone of AI strategy from here.

Implications for Japanese companies and startups

Finally, let me translate all of this into practice for companies in Japan. Here too I write frankly, as my own view.

The starting point is to decide which layer you will fight in. The contest of training a giant foundation model from scratch, meaning going head to head with the world's top at the very bottom layers of chips and models, is not, for many Japanese companies, a realistic main battlefield, given that the order of magnitude of capital and compute is different. Forcing a fight here tends only to wear you down.

On the other hand, the upper layers still hold plenty of room. The application and agent layers that dig deep into a specific industry's work, and a layer that combines your own first-party data with domain knowledge. These can never be filled by a general-purpose giant model alone. The reason is that the tacit knowledge specific to that industry, and the on-the-ground data that only your company holds, are not lying around anywhere on the web. And as we saw in the earlier chapter, the scarcer data becomes, the more the value of this "data only you hold" rises in relative terms. Here, I believe, lies a realistic winning path along which Japanese companies can build a durable position.

This idea of "differentiating through your own data and business processes" is exactly the question of how to build a moat in the application layer. On how to build an application's moat, including network effects, I dig in at my article discussing app moats and network effects, and reading it alongside this piece should give you a more three-dimensional understanding. How Startups Beat Incumbents, which organizes where to build a durable position through strategic frames like the 7 Powers and Lanchester strategy, reinforces this winning path from another angle.

As a concrete move, using an open-weight model as a base becomes a strong option with limited resources. Now that near-world-class models are being published, and cost-effective options are increasing, you can build your own refinements on top of a public model rather than building one from zero. The fact that open-weight models like Kimi have come up to world-class level is itself a tailwind for latecomers. Adaptation on your own data, operation on your own infrastructure, and deep embedding into your business processes. Make the difference with these three. This is, I believe, the realistic path to a winning position without going head to head with vast capital.

Of course, which layer to stand on, and which model to run, where, and in what cost structure, has a different answer for every business. Get the judgment wrong here, and you can wear yourself down in a layer you did not need to fight in, or conversely leak the very data you should be guarding to the outside. If you are hesitating over the design of your own company's AI strategy, please talk to the WARP team. Specialists who led DX and data strategy at major companies walk alongside you month by month, helping to translate the question of which layer, and how to build a durable position in it, into the one step right in front of you.

To sum up

It ran long, so let me organise the key points.

  • The view that "the leading frontier models take everything and you lose" rests on the assumption that AI's strength is decided almost entirely by the volume of data and compute. That assumption itself is, I believe, being rewritten right now
  • The latecomer Moonshot AI (Kimi) combined an open-weights strategy, efficiency in training and inference, a cost advantage, an agent-first design, and fast iterative development to claw its way up to world-class level in a short time. The latest Kimi K3 broke into the world's top three on the Intelligence Index at debut, and is reported to hold a top-tier position on cost-effectiveness
  • That Kimi K3's cost-effectiveness is top-tier is a fact backed by benchmarks. On top of that, my firsthand impression from actually using it is that it is world-class (this assessment is based on my own personal firsthand use, and I keep it distinct from the top benchmark score)
  • The limits of scaling laws are starting to show. With the data wall, as high-quality, human-generated web text nears exhaustion, AI's room to grow is shifting from "gathering data" to "generating new data and experience"
  • In that next stage, synthetic data, reinforcement learning, test-time compute, and autonomous agents become important. These converge on the direction of "the model itself interacting with an environment to create new data"
  • As a result, the axis of competition shifts, in my reading, from "the largest intelligence" to "who can run agents autonomously, efficiently, and at low cost." Here I believe a cost-effective model stands to catch a structural tailwind
  • On layer strategy, AI is a stack of layers from chips to apps, and which layer you fight in, and how far you integrate vertically, is the fork in the road. Both vertical integration and horizontal expansion have their pros and cons
  • Rather than going head to head on training a giant foundation model, Japanese companies and startups would do better, I believe, to build a winning position in the application and agent layers that dig into an industry's work, and in a layer that combines their own first-party data with domain knowledge. It is the path of building on open weights and differentiating through your own data and business processes

The rules of competition do change. That is exactly why there is great meaning in questioning, right now, "on which assumption is my company fighting." Start by taking stock of which layer your company creates value in, and how much, and on how deeply, it depends on which AI. Before you conclude that losing is inevitable, consider whether you can stand on the side where the rules change. That, I believe, is the first step to surviving the AI era ahead.

References and primary sources

Footnotes

  1. Corporate information on Moonshot AI (Yuezhi Anmian). A Beijing-based AI startup founded in March 2023 by Yang Zhilin and others, counted among China's "AI Six Little Tigers." Its funding includes a round led by Alibaba, and as of March 2026 it was reported to be considering a listing on the Hong Kong Stock Exchange (reporting-based information). Moonshot AI official site https://www.moonshot.ai/ , and Wikipedia "Kimi (AI)" https://en.wikipedia.org/wiki/Kimi_(AI) 2

  2. Announcement and specifications of Kimi K3. Announced 16 July 2026. A 2.8 trillion parameter sparse MoE (896 experts in total, 16 activated per input), natively multimodal, with a one-million-token context window. The previous generations were Kimi K2.6 (20 April 2026), the code-specialised K2.7 Code (June 2026), and the first-generation open-source LLM Kimi K2 (July 2025, an MoE with 1 trillion total / 32 billion active parameters, trained on 15.5 trillion tokens). Moonshot announced that K3's weights would be published by 27 July 2026 (awaiting release as of this writing). Efficiency methods such as Kimi Delta Attention (KDA) and MuonClip are based on the company's technical reports (reported). Moonshot AI official site https://www.moonshot.ai/ , Artificial Analysis model page https://artificialanalysis.ai/models/kimi-k3 2 3 4 5 6 7 8

  3. Performance and cost-effectiveness benchmarks for Kimi K3 (Artificial Analysis, AA). An Intelligence Index score of 57, ranked 7th of 190 models on the detail page at the time of viewing, well above the average for models in the same price band. The blended unit price is about 2.31 dollars per million tokens (input 3.00 dollars / output 15.00 dollars, cache:input:output = 7:2:1). The top-tier Claude Fable 5 is reported at an Index of around 60 and a blended price of about 7.70 dollars per million tokens, positioning K3 as a highly cost-effective model that delivers narrowly trailing intelligence at a far lower unit price (there are also reports that it entered the top three on the Index just after debut, though the rank moves as comparison models are added; there are reports of top positions on the agentic knowledge-work benchmark AA-Briefcase and on web-development benchmarks too, all reporting-based and variable by viewing time). On the other hand, K3's output is verbose; when AA ran its evaluation, its output token count reached about twice the median of other models, its generation speed is slow at around 32.1 tokens per second, and the latency to the first response is long (traits of a reasoning model). The low unit price and the total cost on real tasks must be considered separately. Artificial Analysis https://artificialanalysis.ai/models/kimi-k3 (rank, score, and price may vary by viewing time) 2 3 4 5 6 7 8

  4. Open-weights release of Kimi K2. Released in July 2025 under a modified MIT license, and reported to have become the most-downloaded model on Hugging Face the day after release. Wikipedia "Kimi (AI)" https://en.wikipedia.org/wiki/Kimi_(AI)

  5. Estimate of the data wall. Villalobos, P. et al. (2024) "Will we run out of data? Limits of LLM scaling based on human-generated data" (Epoch AI). Estimates the total stock of effective public text at about 300 trillion tokens (90% confidence interval 100 trillion to 1,000 trillion), and projects that, under compute-optimal training, it will be exhausted at an 80% confidence interval at roughly 2026 to 2032 (centred around 2028). With over-training (repeated learning on the same data) it comes earlier (for example, about 2027 at 5x). Because it became clear that filtering refinement yields about 5x and multi-epoch training yields 2 to 5x more, the exhaustion timing has been pushed back relative to the old estimate (which, as of 2022, projected around 2024). Epoch AI https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data 2 3

  6. Remarks by Ilya Sutskever. In his test-of-time award lecture at NeurIPS 2024, he was reported to have said "pre-training as we know it will end," "peak data (we have reached peak data; there is only one internet)," and "data is the fossil fuel of AI," and to have suggested that the next generation heads in a more agentic, more reasoning direction. The primary source is the video of that lecture, but this article confirms it on a reporting basis via TechCrunch and others (reported). 2 3

  7. The "full stack" and "data center = the new unit of compute (AI factory)" framing of AI / accelerated computing by NVIDIA's Jensen Huang. Based on the public stance he has repeatedly presented. However, the specific layer breakdown shown in the body (chip / system / networking / software / model / application) is an organisation and generalisation by the author as a general AI stack, and does not assert the names or order of layers from any specific primary statement. NVIDIA https://www.nvidia.com/ 2

Considering AI adoption for your organization?

Our DX and data strategy experts will design the optimal AI adoption plan for your business. First consultation is free.

Share this article if you found it useful

シェア

Newsletter

Get the latest AI and DX insights delivered weekly

Your email will only be used for newsletter delivery.

無料ダウンロード資料

おすすめの資料

無料診断ツール

あなたのAIリテラシー、診断してみませんか?

5分で分かるAIリテラシー診断。活用レベルからセキュリティ意識まで、7つの観点で評価します。

Learn More About AIコンサル

Discover the features and case studies for AIコンサル.

Related Articles