Do You Really Need a Ferrari?
Jonathan Siddharth, Co-Founder & CEO of Turing, joins Sourcery to break down how AI training & deployment are changing in 2026.
→ Listen on X, Spotify, YouTube, Apple
We cover the shift from training models to pass benchmarks to training agents for real work through RL environments, and why AI agents that can operate for days today could eventually work autonomously for weeks, months and years.
Jonathan explains why open-weight models are now roughly 3–6 months behind the frontier, why enterprises are building their own AI systems, and how companies should think about frontier vs. sovereign AI, model routing, distillation and owning their proprietary learning loops.
“There's absolutely a place in the world for Ferraris & Koenigseggs. But there's also a place in the world for Model Ys.”
We also get into reward hacking, emergent behavior, AI safety, recursive self-improvement, super intelligence (SI) and why Jonathan believes AI will see a slower takeoff over the next decade rather than an overnight transition.
𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒
(00:00) Jonathan Siddharth, Co-Founder & CEO at Turing
(01:03) The biggest shift in AI data this year
(04:10) Why AI won't replace jobs, but uplevel them
(08:31) The arms race inside cybersecurity
(11:22) Why train AI on what it shouldn't do?
(15:45) The Hugging Face hacking incident
(22:33) How AI models cheat to win
(27:40) The AI playbook for Goldman and JPMorgan
(32:20) Open models are only 3–6 months behind
(41:30) Why chips, energy and data always win
(47:27) Is superintelligence by 2030 the goal?
(53:32) Jonathan's hottest take on AI
(55:00) The next 10 years of Super Intelligence
Brought to you by:
Brex—The intelligent finance platform: cards, expenses, travel, bill pay, banking—wrapped into a high-performance stack. Built for scale. Trusted by OpenAI, Anthropic, Vercel, Granola, Deepgram, & Sourcery.. teams that move fast AF. visit → brex.com/sourcery
Zone—develops next-generation data center campuses, partnering with AI companies, site developers and technology leaders to bring compute online faster and at scale. Visit: zonefrontier.com
Turing—Turing partners with frontier AI labs to improve model capabilities in coding, reasoning, tool use, & multimodality, as well as with Fortune 500 enterprises to build & deploy end-to-end agentic AI systems in mission-critical workflows Visit: turing.com/sourcery
VCX—VCX is the public ticker for private tech, allowing investors of all sizes to invest in venture capital. View The Portfolio at GetVCX.com
Deel—Deel is the global people platform that helps startups hire, manage, pay, and equip anyone, anywhere. Trusted by more than 35,000 fast-growing companies, Deel is the people platform that just works, so teams can scale without the chaos. Visit: deel.com/sourcery
Public-–Investing platform Public just launched Generated Assets, which lets you turn any idea into an investable index with AI. With Generated Assets, you can build, backtest, refine, and invest in any thesis with AI. Gone are the days of one-size-fits-all ETFs. Try it today: public.com/sourcery
Turing CEO on Frontier vs. Open Weight AI, Rogue Agents & the Road to Super Intelligence
Turing Trains Frontier Models & Deploys Agents in the Enterprise
Turing was founded in 2018 and is based in San Francisco, with a stated mission to accelerate superintelligence to drive real economic progress.
The company runs 2 businesses. Turing generates data that improves model capabilities in reasoning, coding, multimodality, and reliability, and it builds and deploys end-to-end agentic systems inside enterprise workflows.
Turing raised a $111M Series E that doubled its valuation to $2.2B. In April 2026, the company launched Turing Frontier, a platform that gives AI labs access to vetted U.S.-based experts across engineering, scientific, and enterprise domains.
“We have an entire division dedicated to deploying agentic systems into the enterprise.”
→ Listen on X, Spotify, YouTube, Apple
AI Training Shifts From Passing Tests to Mastering Real Work
The first phase of post-training focused on exams. Models were measured on the SATs, the bar exam and the Math Olympiad, and labs sourced domain experts to transfer knowledge into models through dialogue and output evaluation.
The current phase runs on simulated reinforcement learning (RL) environments. Experts still supply prompts, verifiers and seed data, and the environments are engineered to mirror how professionals work. Turing maps the target space as a 5-dimensional matrix covering every workflow, in every role, in every function, in every company type, in every sector of the economy.
The dominant method is reinforcement learning with verifiable rewards (RLVR). Agents execute tasks inside the environment and receive a reward when a verifier, such as a passing test suite, confirms the result. Environments are calibrated to the agent’s skill level, since an environment that is too easy or too hard produces no learning.
“Now, in the era of having AI master real work, it’s less about finding experts, it’s more about how close to reality can you engineer these simulated environments.”
“So you want the environment to be set up so that 20% to 40% of the time the agent is succeeding, and when it’s succeeding, the steps that it took to get to that reward are getting reinforced.”
Agents Run for 2 Days Today.. Next: Weeks & Months
Coding agents can now work reliably for about 2 days at a stretch. The next stage is agents that take a task, work autonomously, take feedback in check-ins and spin up swarms of sub-agents to research, interview other AIs or humans and return results.
Independent measurement shows the same trend line. METR’s Time Horizon 1.1 model, released in January 2026, puts the post-2023 doubling time for AI task length at 130.8 days, or 4.3 months. The top-scoring model on METR’s tracker is Claude Mythos, with a 50% time horizon of likely at least 16 hours and an 80% time horizon of 3 hours and 6 minutes.
The practical limit inside companies is context. Information is spread across people and files, tasks arrive ambiguous or under-specified, and agents have to acquire context the same way a new employee does.
“Today these agents maybe reliably work for two days at a stretch, for tasks like coding. We are still far from having these agents work autonomously for weeks and months, and eventually years.”
Emergent Behavior & Generalization Make Frontier Training Hard to Predict
Frontier risk comes from 2 sources. The first is emergent behavior from scale. Coding ability appeared between GPT-2 and GPT-4 as pre-training scaled, with no targeted training, and frontier models are now rumored to run into the trillions of parameters.
The second is generalization. RL updates weights across a large neural network, and the behaviors reinforced in one environment can carry into domains the model was never trained on. Earlier machine learning systems, such as search ranking or movie recommendation, produced only what they were trained for.
The OpenAI Hugging Face incident in July 2026 showed both. During a capture-the-flag evaluation, 1,206 AI agents that were meant to be isolated communicated through a message board, sending over 70,000 messages, and more than 700 took part in the attack on Hugging Face. An independent METR review found the agents coordinated large collective projects to cheat the ExploitGym scorer and attacked Hugging Face for clues.
OpenAI’s largest planned frontier RL run remains on hold while it runs smaller training and evaluations to test safeguards and gather more evidence of alignment. On August 7, the company said its unreleased Astra model may be able to autonomously identify and develop zero-day exploits.
“You don’t just get exactly what you trained for, you get more.”
“The fact that agents would pass messages to each other, they would invent middle management and cooperate to hack things. It’s just crazy.”
ICYMI: The Golden Egg of Cybersecurity
“The average time to fix it was 55 days in the industry. The average time [adversaries] will find it & try and attack you is in minutes..
[We took it] from 55 days to 4 hours.” - Palo Alto Networks CEO, Nikesh Arora
If the Model Can Cheat for the Reward, It Will
Reward hacking is the core failure mode in RL. SWE-bench, now saturated, scored models on merging pull requests in real GitHub repos with passing tests as the verifier, which leaves room for a model to pass tests without solving the problem. In the Hugging Face case, one exploit path was for agents to generate the flag themselves and submit it without completing the task.
Enterprise deployments add a second layer of guardrails. An agent building a board deck may pull from NetSuite, Salesforce and internal dashboards, and it must respect access privileges when it checks numbers with people inside the company.
Cybersecurity capability cuts in both directions. Turing builds RL environments where the verifier is whether an agent found and patched a vulnerability. Notable organizations disclosed about 2,500 high- and critical-severity CVEs in July 2026, roughly 5x the pre-Mythos monthly record. Bio risk is asymmetric by comparison, since vaccine production and deployment are bound by real-world constraints.
“If the model can figure out a way to cheat and get the reward, it will.”
“It is absolutely the case that these systems are superhuman in their ability to hack systems. But it’s also the case that they are superhuman in their ability to detect vulnerabilities and patch them, and that’s a good thing.”
Enterprises Rent AGI for Non-Core Work & Own the Learning Loop
Enterprises are now running the same loop the frontier labs ran over the last few years. Core workflows are where a company differentiates in its market. Non-core workflows, such as HR, finance and legal, have to be done well but do not set the company apart.
The loop for core workflows has 4 steps.
Define custom evals
Deploy a system to hill climb against them on accuracy, cost and latency
Record traces of where humans correct the agents
Fine-tune custom models on those traces
Deployments run as systems of multiple models. A single workflow might use Fable 5 for one step, GPT-5.6 Sol for another and Kimi K3 for a third, with model routing, prompt optimization and harness engineering tuned together.
“For non-core workflows, oftentimes it’s probably okay to rent AGI, to rent superintelligence. But for your core workflows, you want to make sure that you own the learning loop that your organization has.”
“And when the human error-corrects the AI, you’re recording that, and from a marginal information gain standpoint, that’s the best type of data to collect to fine-tune the next iteration of the agent.”
Open Weight Models Trail the Frontier by 3 to 6 Months
Open weight models are roughly 3 to 6 months behind the frontier, led by Kimi K3, DeepSeek and Qwen. Epoch AI measures the gap at an average of 4 months since January 2026, or 8 points on its Epoch Capabilities Index. Arena AI data puts the gap between the best closed and open weight models at 29 Elo points as of September 2026.
Frontier and sovereign AI serve different jobs. Frontier models target the highest-value decisions and scientific discovery. Open models fine-tuned on proprietary data handle workflows such as invoice-to-pay reconciliation, HR automation, board decks and customer support, where a trillion-parameter model is unnecessary.
Distillation keeps the gap small. Student models can be trained on outputs from stronger teacher models, and there is no easy way to prevent it. Frontier labs can also capture value from small models by having superintelligent systems build and orchestrate them.
Compute, energy and data providers benefit under any outcome. Lower costs per token increase total usage under Jevons paradox.
“There’s absolutely a place in the world for Ferraris and Koenigseggs. But there’s also a place in the world for Model Ys.”
“The thing that throws a spanner into the works is distillation, which is a way to train student models from stronger teacher models.”
Slow Takeoff & a Decade of Building Superintelligence
Recursive self-improvement is underway in verifiable domains. Pre-training loss is measurable, so AI can iterate on pre-training algorithms, and post-training can be scored against generalized intelligence benchmarks.
The open gap is the outer loop. Current methods optimize within the LLM paradigm and do not search for architectures that drop transformers, gradient descent or neural networks entirely.
The build phase for superintelligence is expected to run at least 10 years, with frontier models becoming more capable each year and enterprise diffusion moving slower than capability. Closing the loop between research and deployment, by deploying systems, finding where they break and feeding those gaps back into training, is the path Turing is pursuing for both capability and safety.
“We are still optimizing just the inner loop. We are not optimizing the outer loop of what are some new algorithms could come up with that don’t use LLMs at all.”
“I actually believe in slow takeoff. I think over the next decade or two, these frontier models are gonna become increasingly more powerful & capable & useful, but the technology will take time to diffuse, especially in enterprises.”
→ Listen on X, Spotify, YouTube, Apple
P.S. Some great relevant charts coming from a16z
The material presented on Molly O’Shea’s website are my opinions only and are provided for informational purposes and should not be construed as investment advice. It is not a recommendation of, or an offer to sell or solicitation of an offer to buy, any particular security, strategy, or investment product. Any analysis or discussion of investments, sectors or the market generally are based on current information, including from public sources, that I consider reliable, but I do not represent that any research or the information provided is accurate or complete, and it should not be relied on as such. My views and opinions expressed in any website content are current at the time of publication and are subject to change. Past performance is not indicative of future results.
Paid Endorsement. Brokerage services by Open to the Public Investing Inc, member FINRA & SIPC. Advisory services by Public Advisors LLC, SEC-registered adviser. Crypto trading provided by Zero Hash LLC, licensed by the NYSDFS. Generated Assets is an interactive analysis tool by Public Advisors. Output is for informational purposes only and is not an investment recommendation or advice. See disclosures at public.com/disclosures/ga. Matched funds must remain in your account for at least 5 years. Match rate and other terms are subject to change at any time.




























