π§ The Dawn of Self-Evolving AI
The AI landscape just witnessed a paradigm shift. MiniMax, a Chinese AI company founded in 2022, has released its M2.7 model, claiming it can recursively improve itself without human intervention. While Google DeepMind's AlphaEvolve and Andrej Karpathy's AutoResearcher hinted at this capability, MiniMax's M2.7 is the first to demonstrate it at scale. According to the company, this isn't just a product update β it's the "early echoes of self-evolution." Sam Altman has described this stage as the "larval stages of recursive self-improvement," signaling that we are ascending a mountain where AI models will increasingly conduct their own machine learning research.
![]()
π οΈ How the Self-Evolution Works: The Harness & The Pilot
MiniMax built an internal research agent harness using an early checkpoint of M2.7. The analogy is simple: the model is the pilot, and the harness is the Formula 1 car. M2.7 was tasked with building its own harness β the data pipelines, training environments, and memory systems needed to manage experiments.
Initially, it acted as a research assistant: performing literature reviews, analyzing proposed experiments, launching tests, fixing bugs, and monitoring results. It handled 30 to 50% of the reinforcement learning team's workflow. This is not just theory; it's a practical, working system that reduces the workload of human AI researchers by a significant margin.
π¬ Step 2: Recursive Harness Improvement
The true breakthrough came when M2.7 began tracking its own performance. It collected feedback on its tweaks, built evaluations for internal tasks, and started iterating on its own architecture. It rewrote its own tools β skills, recipes, and code β to get better at its job. This is the core of recursive self-improvement: the model becomes the engineer of its own evolution.

π€ Step 3: Autonomous Scaffold Optimization
This is where the magic happens. M2.7 ran over 100 rounds of autonomous optimization with zero human input. It followed the scientific method: hypothesize, experiment, compare against a control group, and conclude. If a change improved performance, it was committed; if not, it reverted.
Key Optimization Example: Temperature Tuning
- The model adjusted its own "temperature" parameter. At default (1.0), a poem about cats is conventional. At maximum (2.0), the output becomes wildly creative β even generating demon summoning instructions. M2.7 found the optimal temperature range for its tasks.
End Result: 30% Improvement on Internal Benchmarks
| Benchmark | M2.7 Score | Comparison |
|---|---|---|
| MLE-bench (OpenAI) | 66.6 | Tied with Gemini 3.1 |
| SWE-Pro | 56.22 | Near Opus 4.6 |
| Vibe-Pro (End-to-End) | 55.6 | Near Opus 4.6 |
| Terminal Bench 2 | 57 | Top-tier open-source |
| SWE Multilingual | 76.5 | Highest among open-source |
According to the MLE-bench results, M2.7 earned 9 gold medals, 5 silver, and 1 bronze. It runs entirely on a single NVIDIA A30 GPU (costing $3,000β$7,000), making frontier-level AI accessible to smaller labs and even individual developers.

π’ The AI-Native Organization: M2.7 as an Employee
MiniMax explicitly states that M2.7 is not just a product β it's part of their org chart. The company is using this model to restructure its operations, accelerating its evolution into an AI-native organization. They also launched OpenRoom, an open-source project where an AI agent with a personality interacts with users' files, calendars, and environments. The code for OpenRoom was largely written by AI itself.
π μ 보 κΈ°μ€μΌ: 2024-05-21
Bottom Line: Self-evolving AI is no longer science fiction. M2.7 proves that a model can improve itself, build its own tools, and even manage a company's workflow β all on consumer-grade hardware. The future of AI is autonomous, recursive, and increasingly human-like in its interactions. If you're building an AI-powered business, this is the blueprint.
