TIMEWELL
Solutions
Free ConsultationContact Us
TIMEWELL

Unleashing organizational potential with AI

ISO/IEC 27001 (ISMS) certification mark (SGS / ISMS-AC)

ISO/IEC 27001:2022 certified Certificate No. JP26/00000255 Scope: Planning, development and operation of SaaS products utilizing AI technology

Services

  • ZEROCK
  • TRAFEED (formerly ZEROCK ExCHECK)
  • TIMEWELL BASE
  • WARP
  • └ WARP 1Day
  • └ WARP NEXT Corporate
  • └ WARP BASIC
  • └ WARP ENTRE
  • └ Alumni Salon
  • └ WARP for Schools
  • AI Consulting
  • ZEROCK Buddy

Company

  • About Us
  • Team
  • Why TIMEWELL
  • News
  • Contact
  • Free Consultation

Content

  • Insights
  • Knowledge Base
  • Case Studies
  • Whitepapers
  • Events
  • Solutions
  • AI Readiness Check
  • ROI Calculator

Legal

  • Privacy Policy
  • Manual Creator Extension
  • WARP Terms of Service
  • WARP NEXT School Rules
  • Legal Notice
  • Security
  • Anti-Social Policy
  • ZEROCK Terms of Service
  • TIMEWELL BASE Terms of Service

Newsletter

Get the latest AI and DX insights delivered weekly

Your email will only be used for newsletter delivery.

© 2026 株式会社TIMEWELL All rights reserved.

Contact Us
HomeColumnsAIコンサルOpenAI Announces Next-Generation AI Models O3 and O3 Mini
AIコンサル

OpenAI Announces Next-Generation AI Models O3 and O3 Mini

Published2026-01-21Ryuta Hamamoto
BusinessConsultingAIEventsSecurity

A practical guide to OpenAI Announces Next-Generation AI Models O3 and O3 Mini. Topics include Business, Consulting, AI.

OpenAI Announces Next-Generation AI Models O3 and O3 Mini
Share

This is Hamamoto from TIMEWELL

This is Hamamoto from TIMEWELL.

OpenAI Announces O3 and O3 Mini on the Final Day of Its 12-Day Event

On December 21 (around 3:00 AM Japan time), on the final day of its 12-day new features and models announcement event, OpenAI unveiled the next-generation AI models O3 and O3 Mini. These models significantly exceed the performance of the preceding O1 model and have achieved remarkable results in programming and mathematics. OpenAI hopes these models will mark the dawn of a new era in artificial intelligence.

This article provides a detailed breakdown of the performance of the next-generation AI models O3 and O3 Mini.

O3 and O3 Mini's Remarkable Performance Record-Breaking Results on the ARC AGI Benchmark What Comes Next Summary

Looking for AI training and consulting?

Learn about WARP training programs and consulting services in our materials.

Book a Free ConsultationDownload Resources

O3 and O3 Mini's Remarkable Performance

The next-generation AI models O3 and O3 Mini have delivered stunning results across a variety of benchmark tests, particularly in programming and mathematics — far surpassing the performance of the predecessor O1 model.

Strong Performance on Software-Style Benchmarks

On the SWE-bench Verified benchmark, which consists of real-world software tasks, O3 achieved approximately 71.7% accuracy — an improvement of 22.8 percentage points over O1, demonstrating a significant leap in software engineering capability. In programming, it scored 2,727 on the Codeforces ELO ranking, surpassing OpenAI's Chief Scientist's score of 2,665 and demonstrating advanced coding ability. In mathematics, O3 achieved a 96.7% correct answer rate on a simulated USA Mathematical Olympiad exam, far exceeding O1's 83.3%.

Furthermore, O3 recorded over 25% accuracy on the Epic AI's Frontier Math Benchmark, currently considered the most difficult mathematics benchmark. This is an impressive result given that other AI models had achieved less than 2%.

O3 Mini likewise demonstrated exceptional performance, delivering performance equal to or better than O1 Mini at significantly lower cost. Both in programming and mathematics, O3 Mini outperformed O1 Mini.

Record-Breaking Results on the ARC AGI Benchmark

O3 also set a new record on the ARC AGI benchmark — a test that AI models had long struggled with. ARC AGI is a benchmark designed to measure AGI (Artificial General Intelligence) — easy for humans but difficult for AI. Until now, humans had averaged around 84% correct, while the best AI scores hovered around 30%.

O3 Surpasses Human-Level Performance on ARC AGI

On the ARC AGI private test set, O3 achieved 75.7% accuracy under low-compute settings, placing first on the public leaderboard. Under high-compute settings, it reached an accuracy rate of 87.5% — more than three times better than previous models and surpassing the human average of 85%.

This marks the first time an AI model has achieved human-level performance on ARC AGI. Greg, a representative of the ARC Prize Foundation, stated that this result is an important milestone toward AGI and expressed anticipation for further collaboration with OpenAI.

Currently, O3 and O3 Mini are not yet publicly available. OpenAI is currently conducting internal safety testing as well as providing access to external researchers to verify safety before proceeding with broader release.

However, early access is available for safety and security researchers. By filling out an application form on OpenAI's website, interested parties can participate in safety testing of O3 and O3 Mini and be among the first to evaluate these next-generation models. (Applications were accepted through January 10.)

Public Release Timeline

OpenAI has announced a plan to release O3 Mini to the general public at the end of January, with O3 to follow shortly after. However, the release schedule is subject to change depending on the results of safety testing.

OpenAI has also published a report on a new safety technology called "Deliberative Alignment." Traditional safety approaches train models by showing examples of safe and unsafe prompts to learn the boundary between acceptable and unacceptable content. However, this new technique leverages the model's reasoning capabilities to more accurately judge the safety of prompts, enabling a better tradeoff between safety and performance — paving the way for AI models that are both safer and more capable.

The next-generation AI models O3 and O3 Mini announced by OpenAI have demonstrated remarkable performance in programming and mathematics, achieving human-level performance on the ARC AGI benchmark — a potential harbinger of a new era in artificial intelligence.

OpenAI is taking careful measures to ensure safety, conducting both internal testing and external researcher evaluations. The company is also working on the new safety technology "Deliberative Alignment," aiming to realize AI models that are both safer and more capable.

Public Launch Dates Subject to Safety Results

O3 Mini is expected to be released publicly at the end of January and O3 shortly after, though timing may change depending on safety test outcomes. These efforts by OpenAI represent an important step in advancing artificial intelligence while ensuring its safety.

Reference: OpenAI Official HP "Day 12 — o3 preview & call for safety researchers"

Related Articles

  • The Reality of a Part-Time Employee Who Worked Full-Time, Took Two Maternity Leaves, and Changed Her View of Work | TIMEWELL
  • Before Paternity Leave — What You Absolutely Must Do to Take Leave Even During a Busy Period
  • Pursuing a Hands-On Architecture Firm: Finding My Own Way as the 5th Generation of a Construction Company | Fujita Construction

This article was produced with the help of AI. A human verified the primary sources and edited the text before publication.

Considering AI adoption for your organization?

Our DX and data strategy experts will design the optimal AI adoption plan for your business. First consultation is free.

Book a Free Consultation
Book a Free Consultation45-minute online sessionDownload ResourcesProduct brochures & whitepapers

Share this article if you found it useful

Share

Newsletter

Get the latest AI and DX insights delivered weekly

Your email will only be used for newsletter delivery.

Free download

China-Related Transactions Export-Control Screening Sheet (fill-in / Export Control Law & Dual-Use Regulations, critical minerals, Control List, 2026)

A fill-in working sheet for companies trading with China: screen a single transaction against China's export-control regime (the Export Control Law and the Dual-Use Items Export Control Regulations), the controls on critical minerals (gallium/germanium/graphite/antimony/tungsten etc./rare earths/helium), and the four counterparty-list systems (Control List, Watch List, Unreliable Entity List, countermeasure lists). A procedure for "what to check before the deal," not a roster of "who is listed." With a plain-language intro, based on MOFCOM announcements. Listing is a regulatory category, not a judgment about any company (including the Japanese firms on the Japan-directed lists); controls change continually, so verify current announcements and consult your officer. Match counterparties using the original simplified-Chinese wording.

Download for free

Related Knowledge Base

Enterprise AI Guide

Solutions

Solve Knowledge Management ChallengesCentralize internal information and quickly access the knowledge you need

Learn More About AIコンサル

Discover the features and case studies for AIコンサル.

Contact UsView AIコンサル Details

Related Articles

Inside Tesla's Robotaxi Launch: The Austin Experience, Economics, and the Road to Scale

Inside Tesla's Robotaxi Launch: The Austin Experience, Economics, and the Road to Scale

On June 22, 2025, Tesla hosted a robotaxi launch event in Austin, Texas, inviting selected influencers and Tesla supporters for exclusive autonomous rides.

2026-01-21
Oppo Find X8 Ultra: The 2025 Dream Phone That Refuses to Compromise

Oppo Find X8 Ultra: The 2025 Dream Phone That Refuses to Compromise

A practical guide to Oppo Find X8 Ultra: The 2025 Dream Phone That Refuses to Compromise. Topics include Business, Consulting, AI.

2026-02-07
Smart Home Device Security: What 19 Billion Connected Devices Mean for Privacy and Risk

Smart Home Device Security: What 19 Billion Connected Devices Mean for Privacy and Risk

With approximately 19 billion smart devices in circulation, the attack surface for home and office environments has expanded dramatically.

2026-02-07
Tesla Business Analysis: Tariff Shock, FSD Progress, and the Musk Factor

Tesla Business Analysis: Tariff Shock, FSD Progress, and the Musk Factor

Tesla Stock: The Future Through a Whirlwind. The global economy and technology frontier are in constant flux. Recent markets have been exactly that — a whirlwind.

2026-02-07
Tesla Energy Business: Robotaxi Launch, EV Market Reality, and Electric Aviation

Tesla Energy Business: Robotaxi Launch, EV Market Reality, and Electric Aviation

Part 1: Tesla's Latest Moves — Robotaxi Launch, FSD Evolution, China Recovery. News about Tesla never stops.

2026-02-07
The Future Where Brain Meets AI: How Brain-Computer Interfaces Work, What They Enable, and What Stands in the Way

The Future Where Brain Meets AI: How Brain-Computer Interfaces Work, What They Enable, and What Stands in the Way

Brain-Computer Interface (BCI) — the technology that directly connects the human brain to a computer — once seemed like science fiction.

2026-01-21