Seller Sessions artwork

Cutting AI Costs Without Cutting Output For Amazon Sellers

Seller Sessions · with Sim · September 9, 2026 · 29 min

Summary

This episode offers Amazon sellers crucial strategies to dramatically cut AI operational costs without sacrificing productivity. It highlights how to manage AI spend for teams, switch models effectively, and utilize platforms like OpenRouter and Kimi K3 to turn multi-pound tasks into penny-cost operations. The discussion reframes AI expenditure as an investment, demonstrating how to maximize ROI and optimize workflow for enhanced efficiency.

Key takeaways

Themes

ai & automationamazon & marketplacessupply chain & operationsfinance & fundraising

Topics covered

ai cost managementai model switching strategiesopenrouter cost hackskimi k3 optimizationai roi for ecommercecontext window managementai team usage patternsalternative ai models

Episode description

Danny McMillan and Sim on model switching, OpenRouter cost hacks and whether AI spend for Amazon teams is actually worth it. A quick solo show this week — Dorian's out sick, Matt's unavailable, so it's just Danny and Sim covering how to control AI spend without gutting output. Sim walks through the real cost pressure of a 25-person team on Claude: five figures a year, uncapped Fable usage, and no way to see per-person burn. Danny counters with the setup he's built to fix exactly that — routing grunt work through OpenRouter and Kimi K3 without preloading tools, which took one job from £4-5 down to 35p. They land on a practical split: Fable 5 for planning, a cheaper model for the build, staged "cascade" plans to keep context windows under control. Sim also shares a genuinely wild same-day case study — a 100-video yoga app build for about £300 — and Danny pushes back on the whole framing: stop looking at AI as a cost line and start looking at what it's generating. Key Topics Picking one AI provider and sticking with it - the cost of switching a trained team, and why Astra being better than Fable isn't reason enough to move Why teams actually burn tokens - heavy browser automation, parallel video work, and vibe-coders scaling far past their job description The OpenRouter cost hack - Kimi K3 without preloaded tools, cutting one job's cost by roughly 90% Cascade planning - staging big builds across multiple sessions to control context window burn Chinese model options - GLM 5.3 and LM Studio AI as free-to-cheap alternatives Grok as a low-friction accessory - personal automation for AI-hesitant team members ROI framing - weighing AI spend against what it actually produces, not just what it costs Timestamps [00:01] Quick solo show - Dorian ill, Matt unavailable, flight to Mallorca later [01:02] Wishing Dorian a speedy recovery, back next month [01:17] Kicking off: token costs, Astra's

Related episodes

Frequently asked about this episode

What does this episode say about ai & automation?
Implement a tiered AI model strategy: use premium models (e.g., Fable 5) for planning and cheaper alternatives (e.g., Kimi K3 via OpenRouter) for execution to significantly reduce costs.
What does this episode say about amazon & marketplaces?
Utilize OpenRouter to intelligently route tasks to cheaper AI models like Kimi K3 (without preloaded tools) to achieve up to 90% cost reduction on specific jobs.
What does this episode say about supply chain & operations?
Adopt "cascade planning" for large AI projects, breaking them into stages across multiple sessions to control context window usage and associated token costs.
What does this episode say about finance & fundraising?
Shift the perspective of AI spend from a cost center to an investment, focusing on the value and output generated rather than just the expenditure.
What does this episode say about ai & automation?
Explore alternative and open-source AI models, including Chinese options like GLM 5.3 or LM Studio AI, as cost-effective solutions for various tasks.

Listen