MunHQ

Complete AI infrastructure, from agent workflows to raw silicon.

Most AI setups fall apart at the seams between models, API keys, and compute. We build the whole stack: the agents doing the work, the proxy keeping costs in check, and the GPUs powering it all under the hood.

The yard

Four tools, supplied together or singly.

agentyard The Team

Autonomous teammates that don’t run rogue.

Give AI agents real jobs, clear schedules, and hard spend caps so you never wake up to an empty bank account. Every customer gets an isolated pod and private database. Your company’s data should never mix with anyone else’s.

agentyard.cc
proxium The Gateway

One clean endpoint for every model you use.

Never hand out real provider API keys again. Issue safe virtual keys, fail over seamlessly when OpenAI or Anthropic suffers an outage, and track exact per-call costs down to the penny with budgets that survive server restarts.

proxium.tech
gpuyard The Compute

Stop paying full price for idle GPUs.

Automatically hunts down the cheapest spot compute across seven providers, deploys vLLM on whatever it wins, and spins machines down the second your workload finishes. Works just as cleanly on hardware you already own.

gpuyard.cc
chat-recall The Memory

Your agents stop forgetting what they already did.

Indexes every Claude Code, Cursor, Codex, Gemini CLI and OpenCode session into one searchable history, then hands it back over MCP so the agent resumes its own past work. It also reports any credential that leaked into a transcript.

chatrecall.dev

The tools

Built for our own work first. Every one has source you can read today.

codeindex Code Intelligence

Your agent stops reading whole files to find one function.

A structural code-intelligence engine across 40+ languages, exposed as an MCP server. The agent asks for the symbol it needs instead of burning its context on the entire file.

github.com/munhq/codeindex
cloud-tools Cloud Cost

Find the money leaking out of your cloud bill.

Reads cost, inventory and waste across AWS, GCP, OVH and Cloudflare, as an MCP server and an HTTP API. Read-only access, so it can look at everything and change nothing.

github.com/munhq/cloud-tools
distil Context Economics

Know what compressing a session actually costs you.

Measures where an agent’s tokens go, and whether rewriting the context pays for the prompt cache it throws away. A crate, an MCP server and a benchmark harness.

github.com/munhq/distil

Tell us what you want to run.

One paragraph about the work you want done. We answer with what it looks like and what it costs to run.

General
hello@munhq.com

Send this and we may email you back at the address above. Nothing else.