TokenMizer
an OpenAI-compatible proxy that turns the conversation into a graph of tasks, decisions, files and errors, replayed as a resume block. https://github.com/Shweta-Mishra-ai/tokenmizer
- Core
- 57% · 25.6 / 45 (mean) · no time
- Full
- 57% · 25.6 / 45 · 45 of 100 measured
- Runs
- 5 · core 25.6–25.6 · the details below are the median run
- Common with gbrain
- 25.6 vs 30.4 on 45
- Chars per answer
- 1,271 · answer after 290
- Write → visible
- —
- Write-back
- 0 of 0 degraded answers came back
- Reached through
- HTTP proxy / MCP
- Model on write
- yes: the chat model answers, an LLM extracts the graph
- Key needed
- LLM provider
- Where it runs
- local
- License
- MIT
- Deterministic
- yes in five builds
Configuration
| Version | TokenMizer main at 3c1a19f (6 Oct 2026); it reports itself as 0.5.4, and the 0.5.4 on PyPI (13 Aug) predates the fix that lets LLM extraction reach the graph |
| Mode | OpenAI-compatible proxy: the set is said to it as a conversation, one turn per memory; read through its HTTP API and its MCP server (tokenmizer-mcp over stdio) |
| Door | resume: the resume_context of GET /api/resume/{session}?level=full (live graph), then the text of the MCP tool why_decision for the question; nothing re-rendered |
| Results | the resume block is capped by max_resume_tokens 400 (the default); why_decision answers only decisions; no cut |
| Model inside | chat through the proxy: gemini-3-flash-preview (Google); extraction: gemini-3.5-flash-lite (graph_checkpoint.extraction_model) |
| Embeddings | none configured (semantic_retrieval auto) |
| Storage | SQLite checkpoints in the build's storage_dir, state_backend memory (the default) |
| Not default | use_llm_extraction true (the shipped config has it off). extraction_model pinned to gemini-3.5-flash-lite: the extraction call is capped at 800 output tokens in the code, and gemini-3-flash-preview spends 640-770 of them thinking, so its JSON is cut off and the LLM pass fails every time (3 of 3 in a direct test; flash-lite: 3 of 3 complete). Provider gemini with the gemini extra (google-genai), other providers' keys removed from the server's environment. live_state not run: TokenMizer has none |
Notes: main at 3c1a19f (the 0.5.4 on PyPI predates the extraction fix); chat gemini-3-flash-preview, extraction pinned to gemini-3.5-flash-lite (the extraction call has 800 output tokens and Gemini 3 Flash spends them thinking); door = resume block + the MCP tool why_decision; no live state. Run 2026-10-06, door resume,
fingerprint 7550917880. Report:
examples/tokenmizer_report/reference.md · All runs: examples/tokenmizer_report/repeats.json
Sections
door | 15.6 of 25 · 5 pass · 3 fail | FAIL What port does the Brain listen on? — missing 8766 FAIL Which version of MailBridge is installed for the mailbox? — STALE value delivered next to the current one: 1.3.0 FAIL What is the owner's phone number? — missing +39 |
cards | not measured | section not measured: no entity cards in this memory |
updates | 0.0 of 10 · 0 pass · 1 fail | FAIL the memory serves a fact it was just told — port 8000 not served within 30 s |
time | not measured | only 0 facts out of 22 carry the event date: for the others the age the model hears is the derivation date, not the fact's |
| live state | not run | |
abstention | 10.0 of 10 · 4 pass · 0 fail | |
file search | not measured | section not measured: no repo configured (--repo) |
graph | not measured | section not measured: no case with evidence (all SKIP or no cases) |