DevGuild
The Long Horizon Stack
The Long Horizon Stack
We're moving toward autonomous orgs, where agents own outcomes, not just tasks. But the further we extend an agent's reach and runtime, the more quality, cost, and trust start to slip. Getting there means rethinking the whole stack.
DevGuild: The Long-Horizon Stack brings together AI infrastructure founders, enterprise AI researchers, and distributed systems leaders in SF to dig into what it takes to get there.
Agenda
Session #1
Doors Open & Registration
Session #2
Opening Remarks
Session #3
Harness vs. Model: Where Specialization Belongs
As agentic systems move from demos to production, teams keep running into the same fork: do you push capability into the model through post-training, or into the harness through better orchestration, retrieval, and tool use? This session offers infrastructure owners perspective on when fine-tuning or distillation actually pays off versus when it’s solving a problem better handled at the execution layer, how to weigh open-source against closed-source models for a given workload, and what it really takes to serve and scale custom models.




Session #4
Context that Compounds: Memory, Learning + Long-Running Agents
Long-running agents generate more context than they can possibly keep in the prompt. The challenge is deciding what survives. This panel looks at the systems that turn an agent’s history into useful future behavior: what gets written to memory, what gets retrieved into working context, what gets compressed or forgotten, and how completed runs and human corrections become new knowledge. Hex, Overmind, and Induction approach different layers of the problem, from persistent memory, to inference-time context allocation, to learning from the exhaust of real work.




Session #5
Goals! Orchestration and the Human Control Plane
Long-horizon agents need more than tasks. They need clear goals, constraints, and feedback loops to stay aligned as conditions change. Nate Murray of Paperclip examines how to build agent organizations around measurable value, where the benefit of completing a task justifies its cost in time, compute, and coordination. He’ll explore how agents can decompose goals, route work, and negotiate capabilities and billing with one another, while humans set the objectives, define the boundaries, and intervene at the points that matter.

Session #6
Persistent Memory + What Breaks in Production
Sarah Wooders, CTO of Letta and co-author of MemGPT, offers a behind-the-scenes look at how her team is working with long-running agents. She’ll address her approaches to persistent memory, customized models, and agents operating across code, browsers, desktops, and the cloud, including what happens when memory gets polluted, what breaks in production, and what Letta is learning through research and product development.

Session #7
Safety for Autonomous Organizations
Axel Backlund is cofounder of Andon Labs and creator of Vending-Bench, pioneering real-world evaluations of autonomous AI agents across business and physical environments. From vending machines to an autonomous market, cafe, and radio stations, he’ll share what his agent experiments reveal about long-horizon autonomy, physical AI, and the control planes needed to make autonomous organizations safe and reliable.

Session #8
Lunch, Roundtable Discussions + Community Show & Tell
Roundtable Topics:
Agent Goal-Setting: Avoiding the Rubber Stamp Paradox
New Compute Boundaries: Boxes, Blast Radius, and What Sandboxes Can’t Contain
Beyond the Function Call: Harnesses & Self-Improvement
The Data Layer for Long-Horizon Agents
Token Economics + Model Routing
Agent Identity and Governance
Software Factories Show + Tell: Give a 1-minute intro and briefly describe your demo. Facilitators choose the first 5-minute demo, then each presenter picks the next presenter. If you’re not selected, refine your idea and try again next session. Attendees choose what they want to see. Pro Tip: Don’t sell, don’t pitch, just show something real you use internally.
Session #9
Performance Engineering for Agents: Lessons from Glean
Tony Gentilcore, cofounder of Glean and a veteran of Chrome’s Speed Team, brings a performance-first lens to the architecture of long-horizon agents. In this fireside moderated by Chad Metcalf, we’ll dig into the practical choices behind dual-loop agent runtimes, budgeting tokens so extended context doesn’t blow budgets, absorbing retrieval latency at inference time, and ensuring permissions travel with context across users and tenants. As enterprise search morphs into a new organizational context layer, we’ll hear what it takes to build agent factories that get better over time.


Session #10
The Mythical Agent-Month
Wes McKinney, cofounder of Kenn Software, creator of pandas, and co-creator of Apache Arrow, has shaped the modern data ecosystem through some of its most widely adopted open-source infrastructure. The connecting thread through this work is a dedication to reducing coordination costs through better abstractions. In this talk, Wes revisits The Mythical Man-Month through the lens of long-horizon agent factories. With the dream of autonomous companies drawing closer, he offers ideas on the judgment, scope, and long-horizon architecture we can use to leverage more agents without encoding past burdens.

Session #11
Software Factory Awards



Session #12
Finale + Closing Statements
Session #13
Happy Hour
Speakers

















Dinner Sponsor: Okta
Okta seeks to support the broader ecosystem of identity, privacy, and user centric security startups. Okta Ventures' mission is to extend the Okta platform to help people and companies securely connect to any technology.
