Blog

What does running agent memory actually cost? Setup, retries and upkeep

The real cost of agent memory is not just a bill. It is hosting, model usage, retries and the time of whoever keeps it running. Here is how to count each.

· 4 min read

Running agent memory costs more than the line item you see first. There are four costs: hosting, model usage, retries and re-processing, and your team's time. A managed service bundles them into one service. A self-hosted tool spreads them across invoices and evenings.

We will not invent figures, because they depend on your volume, your models and your people. What we can do is show you what to count.

The four costs

1. Hosting

Self-hosted memory needs somewhere to run: a server, a database, storage, backups and monitoring. Even small setups need all five. The bill may be modest, but it recurs, and it grows as the amount of stored memory grows.

2. Model and inference usage

Memory is not just storage. Turning a meeting transcript or an email thread into facts means calling a language model, and so does answering questions well. With a self-hosted tool you bring your own API keys and pay the model provider directly. That is flexible, but the bill moves with how much your team talks and writes.

3. Retries and re-processing

This is the cost people forget. Memory pipelines run in the background, and background work fails and tries again. Each retry can be a paid model call.

Public issue trackers for self-hosted memory tools show this pattern. Common reports include:

  • Repair or backfill commands that queue paid work the preview did not show, such as one report on repair queues.[1]
  • Estimates of total cost without a hard spend limit, so a run is authorised by page count, not dollars.
  • Identical retries that repeat the same failing, token-heavy request, as in one report on repeated retries.[2]

These are common problems people report with self-hosted memory layers, and teams of those projects work on them. The lesson for buyers is not "avoid them" but "ask what stops a runaway job". A spend cap and bounded retries are worth asking about for any tool, managed or not.

4. Maintenance time

The most expensive line is usually the least visible: someone's hours. Installs, upgrades, configuration fixes, expired credentials, and the moment ingestion quietly stops and nobody notices for a week. We cover this in shared agent memory without a server-maintenance side project.

If the person doing it is a founder or senior engineer, those hours are the costly ones, because the alternative use of that time is building your product.

A way to compare options

Make a small table for each option you are considering. Fill it with your own numbers.

CostSelf-hostedManaged
HostingYou pay and run itIncluded
Model usageYour keys, your billIncluded
RetriesYour spend, your guardrailsProvider's problem
MaintenanceYour team's hoursProvider's job
PredictabilityMoves with usage and incidentsNothing to track

The last row matters more than it looks. A cost that moves with usage and incidents is hard to forecast. A service that carries all four costs for you leaves nothing to track.

What Nysa takes off your list

Nysa is managed, so you carry none of the four costs. There is nothing to host, no model keys to manage, no retries to watch and no upgrades to run. The service covers the brain, the meeting notetaker, the Gmail and Google Calendar connections and the MCP server your agents use, and setup takes under 5 minutes.

Self-hosting fits teams with strict requirements to run everything on your own infrastructure or to rewrite how memory works. Self-hosted vs managed company brain covers how to decide.

Questions to ask before you pick

  1. What is the worst month on the model bill, and what limits it?
  2. If a background job fails, does it retry forever, back off or stop?
  3. Who is on the hook when ingestion silently stops?
  4. Does what you pay change when the team grows or talks more?
  5. What does it cost to leave?

If a vendor, or your own setup, cannot answer the first two, treat the cost as unbounded until proven otherwise.

FAQ

Is self-hosting agent memory cheaper?

Not automatically. Hosting and model bills can look small, but retries and maintenance time often matter more. It can be cheaper if you already have spare capacity and engineers who want the work. Count all four cost types before deciding.

Why do retries matter for cost?

Memory systems process transcripts and emails in the background using paid models. If a job fails and retries without limits, each attempt can be billed. Ask any tool what caps retries and spend.

Does Nysa remove these costs?

Yes. Nysa runs hosting, model usage, retries and upkeep as part of the service, so your team carries none of them. Setup takes under 5 minutes.

Sources

  1. GBrain GitHub issue 6042, https://github.com/garrytan/gbrain/issues/6042, checked Oct 6 2026. ↩
  2. Graphiti GitHub issue 1922, https://github.com/getzep/graphiti/issues/1922, checked Oct 6 2026. ↩

Give every agent the whole story.

We're onboarding a small group of teams.

Request early access