0:00
/

Fireworks AI

How to make AI 10x cheaper with your own models: Lin Qiao, Co-Founder and CEO at Fireworks AI

Subscribe to stay ahead of technology trends. Never miss future editions.

Lin Qiao, co-founder and CEO of Fireworks AI, joins us to explain why “a lot of companies will build AI into bankruptcy” this year. She previously led PyTorch at Meta, and her company now powers every major coding company, including Cursor, which it began working with at $2M ARR and has watched grow a thousandfold.

Lin also shares why the frontier labs’ bet on general intelligence leaves most companies without a moat, how customizing models on private data can make AI five to 10 times cheaper, and why Fireworks expects to process 100x its current 40 trillion tokens a day within a year.

In our in-depth episode, we discuss:

  1. Why product-market fit no longer means a durable business in the AI era, and how companies can scale straight into bankruptcy

  2. The framework for owning your own intelligence: put 10–20% of AI spend into training and cut total cost of ownership by four to eight times

  3. Why the product itself is no longer the moat when coding agents make anything easy to copy, and why your private data is

  4. How Lin runs a company with six additional co-founders, doubling headcount every quarter, on a culture of “extreme ownership”

  5. What Lin learned from Jensen Huang about leadership, and why she believes the future belongs to millions of specialized models rather than a few general ones

Watch or listen now across YouTube, Apple Podcasts, Spotify, and X

Download the transcript 👇

NEW ECONOMIES: Lin Qiao (Fireworks)
282KB ∙ PDF file
Download
Download

Timestamps

(0:00) Meet Lin Qiao
(2:17) Why Jensen Loves Fireworks
(4:15) Jensen's Unique Philosophy
(7:33) What "Human Judgement" Really Means
(11:06) When Founders Should Stop Interviewing
(13:10) Why Join Meta Just to Leave and Build
(17:06) Building With 6 Co-Founders
(22:00) Managing Co-Founder Conflict
(24:47) Why Firms Build on Fireworks
(37:21) The Framework for Owning Intelligence
(38:40) How Long It Takes to Build an Open Model
(43:42) Respecting Privacy in Training
(45:57) The Next Moat
(48:53) Are Data Centers & GPUs an Opportunity?
(50:55) Open vs. Closed Model Strategy
(52:34) Ollie Becomes Lin's Chief of Staff
(54:18) How Lin Uses AI
(56:12) 40 Trillion Tokens a Day

Share

Lessons from this episode with Lin

1. Why companies are scaling AI into bankruptcy
In the SaaS era, product-market fit and a durable business meant the same thing, because the cost of running software was low. Lin explains why AI has split those two concepts apart.

  • Once you hit product-market fit in SaaS, you scaled as fast as possible, because your biggest COGS was people.

  • In AI, your customers can love your product while the cost of serving them exceeds your income. “When you scale, you literally scale into bankruptcy.”

  • For startups, this means running out of money before the next round. For enterprises, it means the CFO looks at the cost forecast and refuses to approve the rollout.

[0:33:30 – 0:35:30]


2. The framework for making AI 10x cheaper
Every company Lin talks to asks the same thing: it sounds like an investment, so how do we justify it? Her answer is to think in terms of total cost of ownership.

  • Put 10–20% of your overall AI spend into training your own model, and the remaining 80–90% into inference.

  • Customizing the model cuts inference costs by five to 10 times, and total cost of ownership still falls four to eight times, even after the training investment.

  • Training isn’t one-off: “We literally launch new models every week,” so the model needs to evolve with the product.

[0:37:00 – 0:38:30]


3. Why your data is the only moat left
Coding agents have made products easy to build and easy to copy. Lin argues the real moat is the data a company collects from customers using its product.

  • Every company has unique product taste, judgment and design, but when anyone can build anything, “their uniqueness may not be that unique.”

  • A company’s customer preferences and feedback are its alpha, and that data will never be shared with anyone else.

  • Yet that data isn’t being used to create any intelligence. That’s the gap Fireworks calls specialized intelligence.

[0:32:00 – 0:33:30]


4. “Are you sure you’re going to start a company with seven of you?”
When Lin pitched Benchmark’s Eric Vishria for her first round, he was surprised by the size of her founding team. Most companies have two or three co-founders; Fireworks has seven.

  • The subtext of his question was that big founding teams usually fall apart through drama and founder dynamics.

  • Lin’s answer is that the co-founders are “brutally intellectually honest” with each other and share an engineering background, which makes logic their common language.

  • Four years in, she says they have become each other’s strength and a bigger force than any one of them alone.

[0:21:00 – 0:22:00]


5. The Navy SEAL principle behind Fireworks’ culture
Fireworks doubled from 150 to 300 people in a single quarter, and Lin still interviews everyone. The trait she looks for is “extreme ownership.”

  • The army is extremely hierarchical, and everyone stays in their lane. On a battlefield, those boundaries stop mattering: you take responsibility, make the call and watch each other’s backs.

  • At Fireworks, “no problem is anyone else’s problem.” It doesn’t matter who wrote the code.

  • It isn’t enough to call out a gap; you have to see it through and fix it yourself.

[0:09:00 – 0:10:30]


6. What makes Jensen Huang different from every other CEO
Lin believes leadership is not a privilege but judgment, and says Jensen operates unlike any conventional CEO.

  • She was shocked by how much detail he understands, from capacity allocation to the technical details of every topic.

  • Traditional companies rely on hierarchy to move information up and down the chain; Jensen simply knows everything across Nvidia.

  • She predicts AI will make this style of leadership far more common, because information will flow freely without deep hierarchy.

[0:05:00 – 0:06:30]


7. How Cursor grew 1,000x, and why knowledge work is next
Fireworks started working with Cursor when it was at $2M ARR, and has watched it grow a thousandfold in three years. Lin saw 2025 as the year of coding and this year as the year of knowledge work.

  • Every major coding company now runs on Fireworks.

  • Lin describes knowledge work as a bushy tree with coding as the trunk. Every profession, from dentistry to legal, has its own depth.

  • Each tip of that tree is an AI-native startup building agents for one specific domain, which is why AI applications are now diversifying so quickly.

[0:25:30 – 0:27:30]


8. From 40 trillion tokens a day to 100x
Fireworks processes 40 trillion tokens a day. Asked where that will be a year from now, Lin weighs two opposing forces.

  • Demand will definitely go up.

  • Token efficiency will rise too: smaller models will deliver the same intelligence and need fewer thinking tokens to reach a good result.

  • Even with those compounding effects, she believes reaching 100 times more tokens within a year is possible.

[0:56:00 – 0:57:00]

Where to find and connect with us

Follow Ollie on X: https://x.com/ollieforsyth

Follow Lin on X: https://x.com/lqiao

Visit Fireworks: https://fireworks.ai

Our partner for today’s episode is Harmonic, your go-to startup database: https://harmonic.ai

Previous episodes include

See all previous episodes here 👉

Discussion about this video

User's avatar

Ready for more?