Calculations are calibrated to the
source data.
Memory counts scale with conversation-token volume. Tokens per memory and latency are measured
strategy coefficients. Maximum combines independently observed maxima and is an upper-bound planning case.
Prices are list prices per 1M tokens, updated July 2026. Embeddings use
text-embedding-3-small at $0.02 / 1M. Query serving, storage charges, caching,
discounts, and infrastructure are excluded.
Redis sizing assumes 1,536-dimensional float32 vectors, 25% vector-index overhead,
4 bytes per memory token, 512 bytes of metadata per memory, and 1 KB per session event.
The model is calibrated to 36.9 MiB/Instruct and 69.4 MiB/Remis + Instruct per 1M source tokens.