Commercial pricing, free quota, and billing examples
Memory started commercial billing on August 20, 2026, 10:00 AM (Beijing Time).
Free quota
Free quota granted after commercialization:
| Item | Tiers | Per-tier quota | Total | Validity |
|---|---|---|---|---|
| Memory write (Add) | 4 tiers | 250 calls/tier | 1,000 calls | 3 months, unused quota expires |
| Memory search (Search) | 2 tiers | 2,500 calls/tier | 5,000 calls | 3 months, unused quota expires |
| Memory storage | — | — | 10,000 entries | No expiry |
Add's 4 tiers: Fragment Pro, Fragment Lite, Profile Pro, Profile Lite. Search's 2 tiers: Pro, Lite.
Pricing
| Item | Plan Version | Price | Billing | Notes |
|---|---|---|---|---|
| Write - fragments | Pro | ¥0.03/call | Per call | Includes inference and vectorization |
| Write - fragments | Lite | ¥0.018/call | Per call | Same |
| Write - profiles | Pro | ¥0.03/call | Per call | Includes inference and vectorization |
| Write - profiles | Lite | ¥0.025/call | Per call | Same |
| Search | Pro | ¥0.001/call | Per call | Includes query vectorization and Rerank |
| Search | Lite | ¥0.00002/call | Per call | Includes vectorization, no Rerank |
| Storage | — | ¥0.002/10K entries/hour | By duration | ~¥1.44/10K entries/month, billed hourly |
Pro vs Lite
The core difference is whether Rerank is enabled during retrieval:
| Plan Version | Rerank | Quality | Use case |
|---|---|---|---|
| Pro | Enabled | Higher, reranking model refines results | Accuracy-critical scenarios |
| Lite | Disabled | Standard, skips reranking | High-frequency, cost-sensitive scenarios |
- Add calls: plan version determined by the fragment rule's
plan_version. Defaults to Pro. - Search calls: plan version controlled by request parameter
plan_version, independent of the rule. Defaults to Pro. - When both
plan_versionandenable_rerankare passed,plan_versiontakes precedence. - Updating a rule's
plan_versionaffects new writes only; existing memories are unchanged. - Pre-commercialization rules default to Pro.
Injecting retrieved memories into the prompt increases Token consumption. Fees are based on actual LLM Token usage.
Billing examples
| Scenario | Calculation | Cost |
|---|---|---|
| Pro write 100 fragment calls | 100 × ¥0.03 | ¥3.00 |
| Lite write 100 fragment calls | 100 × ¥0.018 | ¥1.80 |
| Pro search 1,000 calls | 1,000 × ¥0.001 | ¥1.00 |
| Lite search 1,000 calls | 1,000 × ¥0.00002 | ¥0.02 |
| Store 10K entries for 1 month | 1 × ¥0.002 × 720 hours | ¥1.44 |