The Simulation Premium

The initial wave of AI investing was defined by the economics of pure abundance - funding foundational models that could instantly scrape, summarize, and generate creative content at a marginal cost approaching zero. However, in my eyes, as systems look to transition from the Reading Age (generating text) to the Doing Age (taking action), it's running into a shortage of industry-specific knowledge, and that shift is a much harder problem than abundance alone can solve.

Take chip design. By the 2010s, pushing silicon below the 32-nanometer node required engineers to manually arrange billions of logic gates and sift through millions of lines of simulation logs to diagnose failures. Rule-based tools and brute-force engineering got the industry a long way, but eventually hit a wall. The problem wasn't coming up with new designs. It was testing an idea against physical reality before committing to it.

That same constraint now extends across enterprise software, deep tech, and physical automation. A model can generate a plausible design, a plausible product, a plausible plan. In consumer software, we've made peace with that: nobody checks the math on a Netflix recommendation. The greater challenge is determining whether what it produces will endure the demands of the real world - heat, stress, fatigue, years of use. In critical industries, the stakes are sharper - a flawed wafer, a compromised structure, a system already built - not a line of code you can quietly fix.

Call it the Simulation Premium. As generation flourishes, the scarce resource becomes the ability to test ideas against environments that behave like reality - accurately enough to trust before anything is built. The emerging companies that capture the most value will be those that build the most faithful synthetic worlds - environments where ideas in critical spaces can be tested before reality gets the final say.

The First Test? Physics.

An LLM predicts the next most likely token, but tokens alone can't compute thermal warpage or model structural fatigue across ten thousand load cycles. When the cost of a hallucination shifts from a broken line of code to a melted, multi-million-dollar silicon wafer, the market demands something stricter: verifiable constraints.

Instead of relying on merely plausible outputs, emerging companies like Axiomatic AI are building systems that combine AI-generated engineering reasoning with formal proof frameworks such as Lean to verify physical constraints and design logic before simulation or manufacturing. Through tools like Lemma and Axiomatic Measurement, the company is developing agents capable of reading scientific literature, constructing physics-consistent models, and interacting with physical laboratory environments. One concrete example: Axiomatic built a digital twin of a silicon ring-resonator modulator (a component that lets chips send data using pulses of light, faster and more efficiently than electrical wires) from experimental data in published research. It extracted the transmission spectra, fit a physics-based model for coupling and round-trip loss, and optimized the parameters against the measured response - then used those parameters to generate a photonic layout, letting engineers test and refine device designs virtually before committing to physical fabrication.

Vinci approaches the problem from a different angle, treating geometry and physics as a unified data layer so chip designers can rapidly model heat and mechanical stress in complex 3D packages. Its physics-driven AI platform targets one of the most expensive bottlenecks in semiconductor development: validating designs before committing to costly manufacturing cycles.The approaches diverge, but the underlying bet is the same: the companies that win will be the ones who can compress months of physical trial-and-error into minutes of simulation.

Agent Training

Moving from AI that reads and writes to AI that acts requires more than better models. It requires environments where models can prove themselves before they're trusted with anything real. The industry calls these environments rollouts - simulated end-to-end task executions where an agent navigates enterprise software, fails, retries, and earns a reward only on verifiable completion.

Deeptune represents this emerging infrastructure layer. The critical component of the agentic stack isn't just the model, but the synthetic arena where it gets stress-tested against realistic failure conditions before it ever touches a live system. The fidelity of that arena matters as much as the model inside it - an agent trained on a low-resolution simulation of enterprise software will fail the moment it meets the real thing.

Consider the deceptively simple task of autonomously booking a flight. Training an agent on millions of text guides won't teach it what to do when a calendar widget breaks or a UI button moves. In a sandboxed RL environment, the agent attempts the task thousands of times against a live, fluctuating site. Each attempt produces an unambiguous signal: did the booking complete, or did it get stuck? The agent learns by navigating failure in simulation - not by reading about success.

Adversarial Simulation

Applied to security, the same logic gets more urgent. As agents learn to reason across complex systems and chain minor weaknesses into full attack paths, legacy scanning tools and periodic penetration tests can't keep pace. Finding a vulnerability is one thing; knowing whether it's actually exploitable, before an attacker finds it first, is another - and that's the layer continuous adversarial simulation is built to fill.

Emerging companies like RunSybil are automating the human red team's job: reasoning through enterprise infrastructure the way an attacker would, continuously rather than on a quarterly schedule. Take their pitch at RSAC's 2025 Conference Launch Pad. Traditionally, pentesting meant finding a firm, scoping and scheduling an engagement, waiting for consultants, and receiving a report weeks later. RunSybil's agent, Sybil, compresses that: it maps applications, APIs, cloud infrastructure, and authentication flows, tests for vulnerabilities, chains them into real attack paths, and shows the evidence as it finds it - with retesting available on demand.

Likewise to physics engines and agent training arenas, the value is in the fidelity of the environment, not simply the report it generates.

Why This Isn’t Just a Feature

The natural objection is that simulation becomes just another feature inside existing AI platforms, absorbed by hyperscalers who already control compute. That risk is real. If OpenAI, DeepMind, or Nvidia eventually build general-purpose physics engines, enterprise environments, and verification systems into their platforms, the vertical simulation thesis weakens significantly.

The real question isn't whether foundation models get more capable - they inevitably will. It's whether raw capability alone can recreate the domain-specific environments reliable deployment actually requires.

Three structural forces suggest it can't, at least not easily. Regulated industries - healthcare, aerospace, defense, finance - require simulation environments that satisfy strict validation and audit standards; a general model optimized for plausibility isn't built to satisfy a compliance regime. Customization compounds this: real factories, labs, and enterprise systems run on idiosyncratic processes that a universal model has no incentive to encode one customer at a time. And the highest-value applications often demand private, closed-loop systems, since customers in manufacturing, defense, and healthcare can't expose sensitive operational data to an external platform regardless of how capable it is.

A few of our portfolio companies illustrate this dynamic playing out more broadly - not in simulation specifically, but in the deeper pattern that regulated industries reward: deep domain specialists over horizontal, generalist platforms. Thynk Health (which was acquired by Azra AI) manages specialized medical imaging workflows, lung and breast cancer screening programs, and patient tracking. C2 Keep is an inventory management and compliance platform built for pharmacies to track controlled substances. Virtual Peaker handles industry-specific tasks like grid-edge forecasting, load shaping, and behind-the-meter device integration across smart thermostats and EVs - capabilities that standard enterprise software cannot deliver out of the box. The common thread is that each was founded by industry specialists with the domain expertise to navigate highly regulated, complex fields. Applied to simulation, that same domain expertise becomes the moat - because the harder problem isn't generating a synthetic world, but keeping that world calibrated against reality and within constraints a horizontal platform isn't built to serve.

The market's behavior reinforces this thesis: incumbents are acquiring specialized infrastructure and proprietary environments rather than assuming their own models can recreate them. If foundation models could generate these environments natively, the rational move for well-capitalized players would be to build, not buy - yet capital continues to flow toward companies that have accumulated narrow, hard-won simulation capabilities.

This thesis would be proven wrong if foundation models eventually absorb these calibration loops themselves. But if reliability instead requires years of accumulated domain knowledge, simulation becomes the layer that determines who wins.

The Real Moat: Calibration

Simulation only creates value when the artificial world is accurate enough that decisions made inside it hold up in reality. Getting there is the hard part: continuously encoding the edge cases, constraints, failure modes, and domain-specific knowledge that separate a convincing simulation from a useful one.

A semiconductor simulation company doesn't win on visualization; its models have to understand thermal behavior, material constraints, manufacturing limitations, and decades of accumulated engineering knowledge. An agent training platform wins the same way, by capturing the thousands of subtle workflows, exceptions, and failure states that determine whether an agent survives production, not by recreating a software interface.

Put together, the real moat is the calibration layered on top of the data. The strongest platforms run on a closed feedback loop: more simulations surface more failures, which sharpen the environment, which draws more usage, which generates more calibration data. Call it a learning effect rather than a network effect - instead of users connecting to each other, every deployment improves the system's ability to represent reality.

Foundation models may end up powering the intelligence layer, but the simulation layer is what makes that intelligence worth trusting.

What Customers Actually Buy

...is a reduction in the cost and uncertainty of real-world decisions. That value shows up in three forms:

First, faster iteration cycles. A semiconductor company pays to test thousands of designs digitally before committing expensive fabrication resources. A robotics company pays to train agents through millions of simulated interactions before deploying hardware.

Second, reduced failure costs. A cybersecurity company pays to continuously test attack paths before adversaries exploit them. An enterprise deploying AI agents pays to identify workflow failures before those failures affect customers or operations.

Third, compounding intelligence - the same effect described earlier, now expressed as a purchasing decision. Every failed simulation sharpens the environment, and customers are effectively buying into that improvement over time, not just a single test.

Who Actually Wins?

This dynamic explains why simulation will likely be dominated by vertical platforms rather than one universal environment.

The challenge isn't creating a synthetic world - it's defining the rules, constraints, and failure modes that make that world useful. A semiconductor simulator, autonomous vehicle environment, and enterprise software training ground require fundamentally different representations of reality.

As systems become more general, verification gets exponentially harder, which favors companies that pick a domain narrow enough to model deeply but valuable enough to be worth owning. Those companies become the arena itself - standing between AI-generated output and real-world consequence in a way competitors can't simply prompt into existence.

The Simulation Premium isn't a feature. It's the infrastructure layer of the Doing Age: the most valuable thing in an AI-powered world isn't the output, but the environment that proves it.