icon
[
Talk
]

Long-Distance Relationship: You and Your LLM

Most agentic development today is a long-distance relationship. Your coding agent lives in a datacenter on another continent, you send it your codebase a few thousand tokens at a time, and you pay by the message. It works — but it's not the only arrangement worth being in, and the economics underneath it are less settled than they look. Today's per-token prices reflect a land-grab phase that won't last forever; the workloads you architect around them today are the bills you'll be paying in two years.

At umage.ai we run both. Frontier models in the interactive loop where the extra IQ earns its keep. Local open-weight models on our own GPUs for agents that run inside customer environments, for batch jobs where per-token pricing falls apart, and for workloads where the data legally cannot leave the building. The interesting engineering is deciding which workload goes where — and that decision has been moving fast.

Expect a concrete playbook, honest trade-offs, and the batch-AI category most talks skip: scoring thousands of diffs, analysing every package in an ecosystem, summarising a decade of commits before a migration. Things a pair of GPUs running overnight does well and a per-token bill does not.

[
Other talks
]

An exciting lineup of SPEAKERS.