Back to Y Combinator Startup Podcast

Open Models Change The Economics of AI

Y Combinator Startup Podcast

Full Title

Open Models Change The Economics of AI

Summary

The podcast discusses the rise of open-source AI models and their impact on the industry, highlighting cost reduction, customization, and developer experience as key drivers.

The conversation explores the challenges and opportunities in building and deploying these models, emphasizing the role of platforms like Olama in simplifying access and usage for both individuals and enterprises.

Key Points

  • Open models are significantly reducing the cost of AI for businesses, addressing a major pain point and enabling greater control and customization for specific use cases.
  • There is a cyclical resurgence in interest in fine-tuning custom models, with the rapid release of new open-source models making it challenging to keep pace, though tooling improvements are helping.
  • Chinese-origin open models are seeing significant adoption globally, particularly for AI assistants and coding agents, with countries like Germany being major access points.
  • AT&T is an example of a large enterprise successfully shifting a substantial portion of its token consumption to open models, primarily for coding agents.
  • The increasing context window sizes of open models, now exceeding a million tokens, have fueled an explosive growth in their usage, leading to a massive surge in demand.
  • While custom-trained open models were common, the trend is shifting towards serving out-of-the-box open models, taking off significantly this year.
  • AI safety concerns are emerging as a significant factor, potentially slowing down frontier model development while open-source models continue to advance.
  • Open models offer advantages in specific areas like security testing, where proprietary models may refuse to perform tasks like penetration testing.
  • Olama's role as a key distribution channel for open models involves close coordination with model developers for launches and managing rapid growth spikes.
  • The process of launching a new model involves complex integration with inference engines, harness development, and hardware optimization, often happening under tight deadlines.
  • The developer experience is paramount, requiring seamless integration across various layers of the AI stack, from hardware to applications.
  • The AI landscape is seeing a potential unbundling of services, with specialized companies emerging for knowledge, coordination, and execution, mirroring trends seen in cloud computing.
  • The future of AI likely involves a blend of frontier closed models for cutting-edge tasks and open models for the majority of use cases, with open models dominating token consumption.
  • Local model execution is becoming increasingly viable due to advancements in hardware, offering a cost-effective alternative for certain tasks, though complex coding agents still benefit from cloud models.
  • NVIDIA's strategy of open-sourcing and supporting the ecosystem is crucial for their hardware dominance, enabling a new generation of powerful local AI hardware.
  • The challenges in the GPU market include rapid price fluctuations and supply chain volatility, making it difficult for startups to secure necessary resources.
  • The emergence of ultra-low-cost, high-efficiency open models like DeepSeek Flash is a significant development, enabling widespread adoption and complex task orchestration.
  • The AI development journey for Olama involved pivoting from earlier ideas to focus on open models, driven by a desire to solve developer pain points and a recognition of market shifts.
  • The success of Olama, despite an initial lack of a clear monetization path, highlights the importance of solving a core user problem and iterating rapidly.
  • Y Combinator's community and network proved invaluable for Olama's founders, providing support, lessons learned, and a shared experience during the challenging early stages.
  • The rapid adoption of open models by enterprises and the shift from hobbyist use to mission-critical applications demonstrate a significant acceleration in AI integration compared to previous technological shifts.
  • The underlying principles of building robust software and understanding customer needs, learned from previous tech cycles, are still highly relevant in the evolving AI landscape.

Conclusion

The rapid advancement and adoption of open-source AI models are democratizing access to powerful AI capabilities, driving down costs, and fostering innovation.

Platforms like Olama are crucial in simplifying the complexity of the AI ecosystem, providing a seamless developer experience and enabling widespread adoption by both individuals and enterprises.

The future of AI will likely involve a hybrid approach, leveraging the strengths of both open and closed models, with a continued emphasis on customization, efficiency, and robust developer tooling.

Discussion Topics

  • How are organizations balancing the cost-effectiveness of open AI models with the specialized capabilities of frontier closed models in their AI strategies?
  • What are the biggest challenges and opportunities for startups looking to build innovative applications on top of the rapidly evolving open-source AI ecosystem?
  • As AI models become more powerful and accessible, what ethical considerations and regulatory frameworks are most crucial to ensure responsible development and deployment?

Key Terms

Open Models
AI models whose source code, architecture, and often trained weights are publicly available, allowing for free use, modification, and distribution.
Fine-tuning
The process of retraining a pre-trained AI model on a new, smaller dataset to adapt it for a specific task or domain.
Context Window
The amount of text or data that an AI model can consider at one time when generating a response. A larger context window allows the model to understand and process more information.
Coding Agents
AI programs designed to assist with or automate software development tasks, such as writing code, debugging, and generating documentation.
OpenClaw
A project or framework that enables users to automate complex tasks using open-source AI models, making AI accessible to a wider audience.
Hugging Face
A popular platform and community for machine learning, providing tools, datasets, and pre-trained models, including open-source ones.
GPU
Graphics Processing Unit, a specialized electronic circuit designed to rapidly manipulate and alter memory to accelerate the creation of images in a frame buffer intended for output to a display device. GPUs are also crucial for AI training and inference.
DGX Spark
A high-performance computing system from NVIDIA designed for AI workloads, offering significant processing power and memory.
Apple Silicon
Apple's custom-designed processors for their Mac computers, known for their efficiency and performance in various tasks, including AI.
MLX
A machine learning library developed by Apple for its silicon, optimized for running large language models on Apple devices.
Platform as a Service (PaaS)
A cloud computing model where a third-party provider delivers hardware and software tools—usually those needed for application development—to users over the internet.
Docker Desktop
A cross-platform application that makes it easy to build, share, and run containerized applications.
Kubernetes
An open-source system for automating deployment, scaling, and management of containerized applications.

Timeline

00:00:06

Open models are solving the cost pain point for AI, enabling businesses to gain control and customize AI for their specific needs.

00:00:26

Interest in fine-tuning custom AI models experienced a dip but is now resurging, with the rapid pace of model releases posing a challenge.

00:01:20

The biggest trend in AI is the shift towards open models, driven by coding agents and AI assistants, with models originating from both the US and China.

00:01:51

Chinese-origin open models are being accessed globally through Olama's cloud, with significant usage from the US and Germany.

00:02:15

Cost is a primary driver for enterprise adoption of open models, but the desire for better control and customization is the ultimate goal.

00:02:40

AT&T has transitioned 40% of its token consumption to open models, primarily for coding agent workflows, and is evaluating Chinese models.

00:03:37

Token usage per developer on Olama's cloud shows significant inflection points driven by coding agents and the rise of OpenClaw.

00:04:37

The explosion in open model usage is enabled by improved context windows, now exceeding one million tokens.

00:05:13

While custom-trained open models were prevalent, the serving of out-of-the-box open models has recently gained significant traction.

00:05:54

The increasing speed of open-source model releases makes post-training and customization more challenging, though tooling is improving.

00:06:31

AI safety concerns are becoming more prominent, which may lead to a slowdown in frontier model development while open-source models continue to improve.

00:06:52

The GLM-53 model's cybersecurity capabilities are impressive, presenting opportunities for security and governance startups.

00:07:10

Security and safety concerns are the main blockers for enterprise adoption of open models, but overcoming these can open doors to Chinese-origin models.

00:07:28

Open-weight models are being used to detect hacks, highlighting their utility in security testing, a task proprietary models may decline.

00:08:11

Out-of-the-box open models have some safety training but can better discern between good and malicious use cases.

00:08:26

Olama's role as a distribution channel allows model developers to coordinate launches, leading to significant growth spikes when new models are released.

00:08:53

The rapid release of new models necessitates a well-developed playbook for successful day-zero launches, involving support, performance optimization, and use case identification.

00:09:37

Packaging models with appropriate harnesses and ensuring cloud availability is critical for handling the immense demand upon release.

00:10:03

Open-source harnesses like Codex and OpenCode are crucial for developers to effectively utilize new models.

00:10:35

Hardware and provider collaboration, including optimization for platforms like Apple Silicon, is essential for delivering a good user experience.

00:11:30

The tight integration of drivers, hardware, and applications is key to creating a seamless experience for developers.

00:12:07

Olama's role is to provide a common runtime that matches any harness to any model, simplifying the developer workflow.

00:12:22

Olama's deep technical expertise spans from hardware optimization to complex software stacks, drawing from backgrounds in AI research and distributed systems.

00:12:54

The ultimate focus is on the developer experience, ensuring that API calls return tokens efficiently and reliably.

00:13:09

A challenge with open models has been the lack of coordinated development across the stack compared to frontier model labs.

00:13:42

The "five-layer cake" model of AI development is evolving, with more opportunities for innovation in the layers between models and applications.

00:14:00

The unbundling of services in the AI space, similar to cloud computing, is likely to lead to specialized companies for knowledge, coordination, and execution.

00:14:44

Open source principles are fostering best-of-breed solutions for different aspects of AI, contrasting with bundled offerings.

00:15:05

Frontier model labs aim for walled gardens, while open-source developers seek more flexibility and open ecosystems.

00:15:25

An abundance of open model tokens is being created, shifting the scarcity to orchestration and complex engineering problems.

00:15:53

As coding agents improve, the traditional moat of well-maintained software may diminish, leading to more accessible and interchangeable AI components.

00:16:39

There is a move towards open standards and protocols, reducing vendor lock-in and enabling interoperability.

00:16:48

Core model functionalities like memory might be integrated into models, while staple problems like storage will likely remain separate.

00:17:30

Open models lack the out-of-the-box safety tooling of closed-model providers, presenting a challenge for businesses.

00:17:40

The future enterprise landscape will likely see a balanced allocation of budget between frontier closed models and open-source solutions.

00:18:18

The majority of token consumption is expected to be from open models, though closed models will remain relevant for specific use cases.

00:18:33

Open models are driving down costs, enabling widespread adoption and the development of new orchestration layers.

00:18:48

The collaboration between open and closed models, facilitated by platforms like Open Router, is creating new possibilities.

00:19:47

The analogy of a law firm with partners and associates illustrates how different AI models can work together.

00:20:03

A blend of proprietary and open-source software, as seen in cloud computing, is a common and successful pattern.

00:20:41

Local models are becoming increasingly powerful, with advancements in hardware making it feasible to run larger parameter models on personal devices.

00:21:38

Benchmarks suggest that some local open models are now competitive with top-tier cloud models for specific tasks like coding.

00:21:51

A mix of US and Chinese-trained models are used for local execution, while cloud-hosted models for coding agents are predominantly Chinese.

00:23:14

The stark difference in model origin consumption between local (balanced) and cloud-hosted (Chinese dominant) models suggests a need for more US-based large model development.

00:23:33

NVIDIA's open-source approach and hardware support are crucial for the growth of the open-source AI ecosystem.

00:24:00

NVIDIA's DGX Station and DGX Spark offer powerful on-desk computing solutions for running large AI models at competitive price points.

00:25:22

Next-generation hardware like the DGX Spark, with its high unified memory, enables the running of large parameter models locally.

00:26:02

Both Apple Silicon and NVIDIA hardware are driving a renaissance in personal desktop computing for AI workloads.

00:26:43

Olama's journey started with local models and is now expanding to the cloud, with the expectation that powerful local AI capabilities will return to desktops.

00:27:23

The GPU market is characterized by rapid price changes and supply-demand volatility, posing challenges for startups.

00:27:53

For startups with limited budgets, adopting ultra-low-cost, high-efficiency models like DeepSeek Flash is recommended for maximizing token usage.

00:29:49

Open models have bridged the gap with frontier closed models in terms of intelligence and are now focusing on achieving extreme efficiency.

00:30:02

The DeepSeek Flash model is leading the charge in price-effectiveness and adoption, enabling widespread use of open models.

00:31:01

Orchestrating multiple smaller, efficient models can yield results comparable to larger models, offering a more repeatable and trustworthy approach.

00:31:56

The idea of a single "god model" for all AI tasks has not materialized; instead, a combination of specialized and simpler models, along with orchestration, is proving more effective.

00:32:29

For most use cases, models that are "good enough" combined with architectural improvements are sufficient, though the most powerful models will continue to unlock new possibilities.

00:33:07

Open models are now capable of handling difficult, though not necessarily the hardest, problems, enabling wider adoption.

00:33:13

The possibility of an open-source "god-tier" model exists, with developments in models like ZPoAI and GLM showing frontier capabilities.

00:33:37

The competition between open and closed models is intensifying, moving from a gap to head-to-head parity in certain areas.

00:33:41

Geopolitical considerations in AI often revolve around model origin and where the models are run, with some customers prioritizing US-trained models.

00:34:39

The origin of AI models matters for mission-critical tasks, as it influences the data used for training and the model's behavior.

00:35:07

Concerns about Chinese models, even if hosted in the US, being "booby-trapped" are raised, highlighting the importance of understanding model origins.

00:35:42

Robust IT and security teams in Fortune 500 companies are accustomed to managing dependencies and security risks, as is common with open-source software.

00:36:03

While models are less transparent than traditional software, proper screening and safety checks can mitigate risks.

00:36:14

Olama's origins in the Docker ecosystem and focus on developer experience have led to its strong brand and market position.

00:36:46

Olama's founders, with previous experience building Docker Desktop, spent years searching for the right problem to solve with their developer tools expertise.

00:37:13

Olama pivoted from an initial focus on Kubernetes security to a broader developer security platform before finding its niche with local LLMs.

00:38:39

The name "Olama" was chosen for its character and association with open models (LLM + animal mascot), not directly from a specific model.

00:39:03

Olama's Series A pitch emphasized building a great developer experience and solving security problems, leveraging the founders' previous relationship with investors.

00:39:59

The explosive growth of Olama, beginning in early 2026 according to cloud token data, reflects the rapid adoption of open models.

00:40:41

The initial years of searching for product-market fit were challenging, but the focus remained on solving real customer problems.

00:41:45

A significant pivot involved shifting from security for Kubernetes to developer security on the desktop, which then informed the approach to local LLMs.

00:42:34

The crucial pivot to locally hosted LLMs was driven by a realization of the difficulty in running open-source LLMs and the opportunity to create a seamless gateway.

00:43:48

A bias towards action and rapid iteration led to the launch of Olama within two weeks, quickly gaining significant user traction.

00:44:04

Olama gained initial traction through organic discovery on Reddit, where users praised its ease of use for running local models.

00:45:43

The transition from hobbyist users to 85% of the Fortune 500 in a short period mirrors the speed of PC adoption after the homebrew computer club era.

00:46:05

The ease of getting started with open models, coupled with their ability to run anywhere, was a key factor in their adoption by enterprise developers.

00:47:04

Olama's monetization journey began with a privacy-focused AI product and evolved as open models became capable of handling complex tasks like coding agents.

00:47:40

The monetization strategy for open models involves offering privacy-focused AI products and capturing value as these models mature and handle more demanding use cases.

00:48:57

Waiting for the market to mature allowed Olama to identify and address the most critical problems for customers using open models for challenging tasks.

00:49:40

Y Combinator's support was crucial for Olama, providing a peer network and lessons learned from previous generations of startups.

00:50:56

The AI world presents new challenges and opportunities that differ from traditional infrastructure, requiring new lessons learned.

00:51:36

The experience of working on teams that have shipped successful technology with clear product-market fit provides valuable muscle memory for building new ventures.

00:53:37

The AI landscape is breaking some traditional DevOps and infrastructure rules, necessitating new approaches to building and scaling services.

00:55:07

Olama provides a curated and seamless experience for end-users, abstracting away the complexities of underlying inference providers and their potential issues.

00:56:07

Curation of the fragmented open-source AI ecosystem, including models, inference technology, and services, is a valuable contribution to developers.

00:56:40

Platforms like Open Router and OpenCode simplify model selection, payment, and integration, empowering developers to focus on building applications.

00:56:53

In an abundance of open models, scarcity shifts to bringing together these fragmented components into a cohesive and functional solution.

Episode Details

Podcast
Y Combinator Startup Podcast
Episode
Open Models Change The Economics of AI
Published
September 12, 2026