Skip to main content

3 posts tagged with "infrastructure"

View All Tags

Open Sourcing Moyai: Self-Hosted Cloud Coding Agent

Ishaan Jaffer
CTO, LiteLLM
Tin Lo
AI Product Engineer, LiteLLM
Moe Khalil
AI Product Engineer, LiteLLM

MoyaiOpen source cloud agent

October 7, 2026
HermesHermes
Claude CodeClaude Code
CodexCodex
OpenCodeOpenCode
Deep AgentsDeep Agents
OpenAIOpenAI
AnthropicAnthropic
FireworksFireworks
GoogleGoogle
xAIxAI
MistralMistral
DeepSeekDeepSeek
BedrockBedrock

Works with Claude Code and CodexSelf-hosted · 100+ providers through LiteLLM

  1. [1]The problem
  2. [2]Our cost estimate
  3. [3]Why we're open sourcing it
  4. [4]A cloud agent that keeps working
  5. [5]Choose your agent
  6. [6]Choose your models
  7. [7]Know what you spend
  8. [8]Set up your first task
Published:
Ishaan JafferCTO, LiteLLMTin LoAI Product Engineer, LiteLLMMoe KhalilAI Product Engineer, LiteLLM

Today we're open sourcing Moyai, the cloud coding agent our team runs at LiteLLM. Give it a task in Slack or the browser, then come back to a pull request with code and test results to review. You choose the agent and compatible model, run it on your infrastructure, and track the cost through LiteLLM.

Set up your first cloud task, or see the cost comparison and a real bug fix below.

How we built our own internal Devin in 2 days

Tin Lo
AI Product Engineer, LiteLLM

How we built our own internal Devin in 2 days

Our team's cloud coding agent, built with Render and Temporal.

Our internal Devin, running in the cloud.
Published:
Tin LoAI Product Engineer, LiteLLM

At BerriAI, we built Moyai Devin, an internal engineering agent that runs in the cloud. Teammates can give it a task in Slack, follow its progress in a web app, and ask it to prepare a pull request.

Making the AI Gateway Resilient to Redis Failures

Ishaan Jaffer
CTO, LiteLLM

Last Updated: April 2026

Enterprise AI Gateway deployments put Redis in the hot path for nearly every request: rate limiting, cache lookups, spend tracking. When Redis is healthy, the latency contribution is single-digit milliseconds, invisible to end users. When it degrades, a production AI Gateway needs to stay up regardless.

Running LiteLLM at scale across 100+ pods means designing for failure modes before they appear. The easy case is Redis going fully down: fail fast, fall through to the database, continue serving requests. The hard case, the one that takes down gateways, is a slow Redis: still accepting connections, still responding, but timing out after 20-30 seconds per operation.