Hello, this is Ryuta Hamamoto from TIMEWELL.
The most common sentence I hear about AI adoption: "We ran a PoC, but it hasn't gone anywhere since."
Someone tried it. It worked. There was a report. And that was that. The deck from six months ago is still in the shared drive and nobody has opened it.
This piece is written from the point of view of the organisation buying AI. Why pilots stall, and what to settle up front so they do not. It is about decisions, not technology.
The short version:
- They stall not from enthusiasm or technology but because nobody set the criteria for going to production
- A PoC confirms whether it works. Production needs an owner, a cost, and maintenance
- Japanese AI usage is concentrated in efficiency. 91.6% on efficiency against 3.9% on revenue
- Run only as many pilots as you can shepherd all the way through
The numbers say this is structural, not local
Start with evidence that this is not your company's private problem.
IPA published a survey in July 2026 covering 1,799 Japanese companies, fielded between April and June1. The breakdown of what AI actually delivered is stark.
| Reported effect | Share |
|---|---|
| Work became more efficient or faster | 91.6% |
| Quality and speed of proposals improved | 48.9% |
| Overtime reduced | 29.2% |
| Customer satisfaction improved | 4.5% |
| Revenue or profit improved | 3.9% |
| Customer base expanded | 2.7% |
Nine in ten feel the efficiency. Fewer than four in a hundred see it reach revenue.
The usage data explains why. Summarising, translating and proofreading text or audio sits at 82.5%; drafting documents and reports at 80.5%; search, collection, analysis and reporting at 77.0%. Meanwhile, upgrading the company's own products and services is 10.9%, and planning support for production, logistics or service delivery is 6.0%1.
Summarising and drafting. That is a sensible way in — quick returns, low risk. The problem is that there is no design for going past it, and the whole country has stopped there.
That shows up directly in how pilots get scoped. You confirm that drafting is faster, and it is. Then the question "so what does production mean here?" arrives, and nobody has prepared an answer.
The cause is an undefined exit
Here are the patterns I see. None of them is a technical failure.
1. No definition of what justifies proceeding. The most common by far. Start with "let's just try it" and there is nothing to judge at the end. It worked — and? With no criteria, the conclusion becomes "let's keep an eye on it."
2. No production owner. Pilots usually run out of IT or a strategy function. Production runs in the business. Who operates it, who fields questions, who notices when accuracy degrades? You cannot discuss production with those blank.
3. The money comes from a different pocket. Pilot budget and operating budget are different lines. Different lines mean rebuilding the business case from scratch. Put a rough production cost in the pilot proposal or it will not even reach the right approver.
4. Success is defined as accuracy. 95% passes, 90% fails. Set it that way and you cannot judge whether the thing is usable. Some workflows run fine at 85% with a human check; some need more than 99% because safety is involved.
5. Too many in parallel. Pilots are easy to start. One per department and you have five in a month. None reaches production and the coordinator's capacity quietly evaporates.
The common thread is that things which should be settled before the pilot are being deferred until after it.
Looking for AI training and consulting?
Learn about WARP training programs and consulting services in our materials.
Five things to settle up front
One. Which number, in which process, moves by how much.
Not "try generative AI" but "reduce the time to produce one quotation from 40 minutes to 20." If you do not know the current figure, measuring it is step one of the pilot. Skip that and you cannot claim an effect afterwards.
Two. The condition for going to production.
"If the pilot demonstrates X, production is approved." Agree that sentence with the decision-maker before starting. Put the other way: a pilot where you cannot write that sentence is not ready to start — it means what you want to learn is not yet defined.
Three. Who operates it in production.
Which department owns it, who runs it day to day, where questions go. "We'll decide that when it goes live" is a reason it never goes live.
Four. The cost picture.
Not the pilot cost — a rough annual production cost. Licences, infrastructure, operating labour, model updates. If this is out by an order of magnitude, a successful pilot dies on "too expensive." It does not have to be precise. It has to be the right order of magnitude.
Five. The condition for stopping.
Few companies set this. Decide in advance what result means you do not proceed. Without it, a pilot that is not working continues indefinitely on "let's try a bit more." With a stopping rule, you get a conclusion and move on to the next topic.
Write those five on one page before you start, and most PoC limbo disappears.
Design what happens after the efficiency
Back to the opening numbers. 91.6% efficiency, 3.9% revenue.
Closing that gap does not need a better model. It needs a decision about what the freed-up time is for.
Say quotations go from 40 minutes to 20. Five a day is 100 minutes a day, roughly 35 hours a month. Can you say what those 35 hours are being spent on?
If not, the efficiency will not become a number. Work got slightly easier, overtime dropped a little. Both worth having, and neither appears in revenue. The IPA data is a picture of exactly that state.
To reach revenue you have to decide the destination in advance. "Halve quotation time and use it to add ten customer visits a month" is where it becomes a revenue conversation. Write that and the pilot's metrics change too: not just minutes per quotation, but visits and win rate.
This is not a technology problem. It is the buyer's job to articulate what they actually want. Leave it unarticulated and the vendor will build something technically correct — something that works and does not move revenue.
Questions the buyer should be asking
There are questions to put to anyone you are considering. Put differently: pick the party that can answer them.
"If this pilot succeeds, what happens next?" No concrete answer means the proposal does not extend to production.
"Who operates it, and how?" This tells you whether it is a build-and-leave proposal.
"If it does not work, where do we stop?" Anyone who can discuss stopping is usually worth trusting.
"Who inside our company will be able to keep this going?" That is the insourcing question. Hand everything over and the next change is another procurement. There is more on where to draw that line in how far to insource AI.
When we take on AI adoption work through WARP, the first step is not tool selection. We start by asking what currently takes time, and what you would do with that time if it were freed. Pick the tool before settling that and the pilot usually stalls.
Pilots are still worth running
None of this is an argument against pilots. Trying something small beats rolling out company-wide and failing. The problem is when trying becomes the objective.
And one more thing. "It stopped at PoC" is not necessarily a failure. If you tried it and concluded it should not go to production, that is a correct decision.
The real failure is no decision at all. Not a success, not a failure, just over. When the same topic comes round again, nobody can say what was learned last time. That is the expensive outcome.
Keep the records. What was tried, what was learned, why it did not go to production. With that on file, the next pilot starts where the last one ended.
In summary
- Pilots stall because the criteria for production were never set
- Japanese AI usage concentrates on efficiency. 91.6% efficiency against 3.9% revenue
- Settle five things first: the number that moves, the production condition, the owner, the cost, the stopping rule
- Decide where the freed-up time goes or it never reaches revenue
- Ask "if this succeeds, what happens next?" before you commit
- The real failure is a pilot nobody decides on. Keep the records
To talk through how to approach AI adoption, or to design the pilot itself, get in touch.
References
Footnotes
-
Key points on DX and AI adoption trends among Japanese companies (IPA, 16 July 2026, Japanese). The reported effects of AI adoption (91.6% efficiency, 3.9% revenue or profit) and the usage breakdown (82.5% summarising and translation, 10.9% upgrading own products and services) come from this material. Respondents were executives, IT departments and DX functions at Japanese companies; 1,799 responses, fielded 17 April to 12 June 2026 ↩ ↩2



![Why 40% of AI Adoption Projects Fail [2026 Edition]: Three Traps Revealed by Stanford HAI's 88% and Gartner's 40% Cancellation Forecast](/images/columns/ai-adoption-failure-traps-stanford-gartner-2026/cover.png)

