Hugging Face funes: data boundaries and the 8x denominator
funes, released 3 September 2026, indexes coding-agent sessions. Here is where the data sits at each of three stages, and the denominator behind the 8x figure.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Bottom line: local by default, and nothing leaves until you bind
Three points carry most of the practical weight in the funes release of 3 September 2026:
- With no memory bound, indexing stays on your machine. The qualifier matters: a hosted model does not process your sessions for indexing, and your coding agent does the reasoning
- Data reaches the Hub when you bind a memory with
funes add. Repositories funes creates are private by default, but existing repositories keep their current visibility - Making a memory public is irreversible. The remote is append-only, and nothing retracts a session once it is up
Everything below comes from Hugging Face itself. As of 4 September 2026 we found no third-party verification.
What shipped
Hugging Face describes funes as a durable memory layer for coding agents: Claude Code, Codex, pi and Hermes.
| Item | Detail |
|---|---|
| Announced | 3 September 2026, on the Hugging Face blog |
| Code | github.com/huggingface/funes |
| License | Apache License 2.0 |
| Primary language | Rust |
License and language come from the GitHub repository record, retrieved 4 September 2026.
Reading “supports four agents” flatly loses the staging in the repository:
| Agent | Condition stated in docs/add.md |
|---|---|
| Hermes | Per-turn indexing is marked beta (shell hooks) |
| Codex | Publishing at session boundaries needs Codex 0.151.0 |
| Claude / Codex / Hermes | funes add runs the one-time bootstrap (first index, hooks, first push) |
Where the data sits, stage by stage
| Stage | Action | Storage | Who can read it |
|---|---|---|---|
| 1 | Nothing (default) | Local Lance dataset | Your machine only |
| 2 | funes add <agent> <org>/<repo> | A Hugging Face dataset you own | Whoever holds repository access |
| 3 | Flip the repository to public on the Hub | Same dataset, public | Anyone, via the --memory flag |
Stage 1: what runs locally, and what does not
| Step | Where it runs |
|---|---|
| Parsing, chunking, embedding, reranking | Your machine |
| Reasoning | Your coding agent |
Recalled passages enter that agent’s context, so “sessions never leave” overstates the claim. SECURITY.md is conditional: nothing leaves your machine unless you run funes push, or bind a shared memory whose hooks push at session boundaries.
Stage 2: how the default visibility actually works
Repositories created by funes are private by default. Existing ones keep their current visibility, and making a funes-created memory public is a deliberate change on the Hub. SECURITY.md asks the user to verify visibility before pushing history they want kept private.
Stage 3: why it is irreversible
docs/push.md states that the remote is append-only, and that selecting what to publish is a pre-publication gate rather than a remote undo. The published example, huggingface/funes-memory, lists 27,118 chunks and embedding model BAAI/bge-small-en-v1.5 as of 2 September 2026.
What the secret scanning does and does not promise
| Layer | Action | Condition stated in the sources |
|---|---|---|
| Index-time redaction | Removes credentials before a session is stored | docs/push.md scopes it to “when TruffleHog is available” and calls the pass best-effort; without it, a warning is printed and indexing continues |
| Publish-time gate | Scans outgoing chunks, withholds rows that still look like secrets, exits non-zero | Fail-closed: if the scanner is missing or crashes, nothing is published |
funes scrub | Removes secrets from rows already in local memory | Does not alter an already-published remote |
SECURITY.md is explicit that the gate stops future rows and cannot unpublish. Fail-closed means “publish nothing if the scanner cannot run”, not “no secret gets through”.
The 8x and 4x figures, with their denominator
Hugging Face reports that recall cost an eighth of a written handoff on one task and a quarter on the other.
| Aspect | Detail |
|---|---|
| Scale | 5 channels x 2 tasks x 3 reps = 30 runs |
| Tasks | rerank-triage and recall-features only |
| Axis | Weighted tokens per successful task |
| Weighting | input + 1.25 x cache creation + 0.1 x cache read + 5 x output |
| Dollars | Excluded from the results page: billing tier depends on the configured context window |
Results by channel (weighted tokens per successful task)
| Channel | rerank-triage | recall-features |
|---|---|---|
| A Nothing carried over | never arrives | never arrives |
| B Written handoff | 851k | 637k |
| C Recall (funes) | 101k | 169k |
| D Whole session kept alive | 822k | 651k |
| E Compact and continue | 778k | never arrives |
851/101 is 8.43 and 637/169 is 3.77, so the published “4x” rounds up.
The one-time preparation charge decides the ratio
| Channel | How the one-time charge is treated | rerank-triage | recall-features |
|---|---|---|---|
| B (handoff) | Added once in full, not divided across the three reps | 808k | 538k |
| E (compaction) | Same | 734k | 549k |
| C (recall) | None recorded (funes index incurs no API spend) | — | — |
Looking only at the production turns of a single rep, B costs 39k, 39k and 50k while C costs 148k, 78k and 78k. Excluding preparation, the handoff is cheaper, and the dataset card concedes that this charge usually decides both comparisons. If the same handoff is reused across many sessions, the published ratio does not carry over.
The source also flags that the 808k charge on rerank-triage is a reconstruction: that receipt was lost and the figure was rebuilt from the run’s printed summary. The 538k has a surviving receipt.
Compaction split across the two tasks
Hugging Face reports that compaction was the only channel whose result divided: it arrived on every rep of rerank-triage and never on recall-features, where the summary flattened the findings that mattered. No cost per success is not the same as no cost: on recall-features it spent 648k to land where the carry-nothing channel lands. The results page limits the conclusion itself, calling this compaction losing this investigation’s findings and not a claim about compaction at large.
Claims to avoid when citing this release
- “Session logs never leave your machine.” The qualifier is for indexing; reasoning happens in your coding agent
- “A shared memory is always private.” Private by default covers repositories funes creates
- “8x cheaper means one eighth of the bill.” The axis is weighted tokens; dollars were removed as not comparable
- “Recall is always 8x cheaper.” The measurements are 8.43x and 3.77x on two tasks, and depend on charging preparation once in full
- “Compaction loses information, so avoid it.” The source forbids that generalization
- “It works the same on non-English sessions.” The source says nothing about other languages
Related
- IBM Research measured agent memory across 8 models: more is not better — how much extracted guidance to inject
- @huggingface/kernels 2.57x: the denominator and test setup — another ratio put back into its conditions
Sources
- Give Your Coding Agents a Memory You Own (Hugging Face Blog)
- huggingface/funes README.md (GitHub)
- huggingface/funes SECURITY.md
- huggingface/funes docs/add.md (supported agents and binding)
- huggingface/funes docs/push.md (publishing and sharing)
- dacorvo/funes-handoff-recall-benchmark results/README.md
- dacorvo/funes-handoff-recall-benchmark dataset card
- huggingface/funes-memory (a published memory)
- GitHub REST API: repository metadata for huggingface/funes
この記事の日本語版: Hugging Face funes: data boundaries and the 8x denominator(日本語)