60 Commits

Author SHA1 Message Date
durmazoguzhan c46b756cde feat(merge): link a type declaration two repos share (#3007)
When merging graphs from multiple repos, a type declared in more than one repo under the
same fully-qualified namespace and name (a shared contract type) now gets a same_type_as
edge linking the declarations, so a cross-repo contract is navigable. Matching requires a
non-empty namespace plus label, a real sourced type declaration (not a method/field or
sourceless stub), and at least two distinct repos, so two unrelated types that merely share
a short name are not linked; the edge is INFERRED/0.9.
2026-08-24 14:09:46 +01:00
SinghAman21 245eb9dc9c fix(extract): scope fresh semantic results to dispatched files (#2926)
The #1757 guard scoped the semantic-cache WRITE to dispatched files, but the unfiltered
fresh result still fed build_merge, whose replace-set logic swapped a non-dispatched
file's entire prior contribution for a stray misattributed fragment (and logged the
'skipped out-of-scope source_file' warning). Apply the same allowlist to the result dict
before it reaches the merge, via a shared _semantic_source_matcher so the write guard and
the graph filter cannot drift.
2026-08-21 19:21:20 +01:00
hopstreax 24cf52d51e fix(cache): retry zero-node semantic results (#2927)
A semantic result with no nodes and no hyperedges (only edges, or nothing) was cached and
stamped into the manifest, so an empty/degenerate LLM reply for a file froze that file:
detect_incremental saw it unchanged and never re-dispatched it. Reject zero-node results
from the cache read and write, drop edges from the manifest stamp tuple, and heal an
existing manifest by re-queueing files that were already stamped with a zero-node result.
2026-08-21 18:26:49 +01:00
santhiprakash 5be36dac0d fix(extract): preserve existing semantic layer on --code-only --force (#2923)
A `--code-only --force` rebuild over an existing graph dropped the doc/paper/image
semantic tier, because force took the full-rebuild path and code-only never re-dispatched
those files. When an existing graph is present, keep incremental mode so build_merge
carries the semantic layer forward; files deleted from disk are still pruned.
2026-08-21 18:26:24 +01:00
rajarshidattapy eccc6cf0bf feat(extract): add --no-dedup so incremental merges can skip fuzzy dedup (#2881)
Adds --no-dedup (default off, so dedup stays on) to skip the fuzzy near-duplicate merge
pass on build and incremental merge, for operators who would rather keep distinct symbols
that fuzzy-matched than pay the merge. Exact-id uniqueness is unaffected (it is a graph
structural invariant, not a dedup responsibility), and the flag arms the #479 shrink
guard so a surprising node drop is refused loudly. Mutually exclusive with --dedup-llm.
2026-08-20 16:10:49 +01:00
oleksii-tumanov 05f90f1682 fix(export): restore graph.html for large graphs (#2853)
The label and cluster-only commands did not pass the viz node limit to to_html, so a
graph over the node limit raised and the except branch silently unlinked the graph.html
that update had produced. Always pass the limit (the aggregated community meta-graph
renders instead of raising), preserve the existing file on a failed render via an atomic
publish plus a stale marker, and regenerate a missing graph.html on the no-topology-change
fast path without reclustering.
2026-08-20 16:10:36 +01:00
abhay-codes07 e4b85d3b50 fix(hooks): a driveless rooted path is not cwd-relative on Windows (#2795)
The out-of-project read guard treated any non-absolute path as cwd-relative, but on
Windows a rooted-but-driveless path (`\foo\bar`) is not absolute yet resolves against
the current drive root, outside the project. Classify with a platform-correct predicate
(`not root and not drive` under the host path flavour) so the containment check is
reached. On POSIX the predicate reduces to `not is_absolute()`, so no behaviour changes
there.
2026-08-19 14:05:58 +01:00
abhay-codes07 b8a8b4c4c9 fix(query): name the graph the answer came from (#2789)
The query header now leads with the graph file it opened and its node count,
shown relative to the CWD when the graph sits underneath it and absolute when it
does not — surfacing the case where a query run from a parent project silently
answers from the wrong corpus. graph_path is optional, so the header is
byte-identical for callers that do not pass it; the CLI query command and the MCP
query tool, which already resolve the path for querylog, now pass it.
2026-08-18 22:36:42 +01:00
Ousama Ben Younes 243a1801f1 fix(affected): resolve an absolute-path seed via the graph-derived root (#2706)
affected anchored an absolute-path seed to Path.cwd(), so running it from
anywhere but the repo root made relative_to(cwd) raise and the query fell
through unmatched, silently returning nothing. It now derives the repo root
from the graph's own location (<root>/graphify-out/graph.json) and anchors the
seed there, so an absolute seed resolves regardless of cwd. Composes with the
#2707 relative-seed fix (root defaults to cwd for other callers).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-15 17:12:16 +01:00
C0KERNEL cc0ee60c40 fix(export,cli): stamp graph provenance from the analysed repo, not the shell cwd (#2534 family) 2026-08-13 13:30:24 +01:00
safishamsi 09a34ad87a fix(update,llm): retry failed extractions; surface claude-cli envelope errors; bump to 0.9.37 (#2543, #2554)
#2543 (adopts PR #2546, thanks @michaelxer): a failed extraction is no
longer stamped in the incremental manifest as up-to-date, so graphify
update retries it instead of skipping it forever; a manifest already
poisoned by the old behavior is healed on the next run; genuinely
unchanged files are not re-processed. Extended to the watch save_manifest
paths too.

#2554 (adopts PR #2555, thanks @annieyii): the claude-cli backend now
inspects the stdout envelope for an is_error result (e.g. a rate limit
returned with exit code 0) and raises it on both the zero and non-zero
exit paths, instead of parsing it as an empty success and bisecting
against a live rate limit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-08 23:37:49 +01:00
safishamsi 6ba0868228 fix(cli): surface four silent success-exit failures (#2534)
cluster-only warns when --backend/--model/--batch-size are ignored on the
label-reuse path; the community-label prompt key no longer collides with
the discard sentinel (an echoed key was silently dropped); tree --root
exits non-zero when it matches no source file instead of flattening the
tree; and cluster-only stamps built_at_commit from the analysed graph, not
the shell cwd. Also folds in the cluster-only refused-write guard from
PR #2522 (thanks @aniJani).

Thanks @elecnix for the report.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-07 22:49:25 +01:00
safishamsi 07b9143d4b fix(hyperedge,skill): merge/load hyperedge integrity + community labels; bump to 0.9.34
#2486 (thanks @adminwat): normalize dict-shaped hyperedge members to ids
(or drop with a warning) so a malformed hyperedge can't abort a completed
merge with a TypeError.
#2484 (thanks @sortakool; approach from @oleksii-tumanov's #1691):
merge-graphs relabels hyperedge member ids and ids with the repo prefix,
unions both inputs' hyperedges instead of clobbering, and writes both
persistence slots.
#2485 (thanks @sortakool): build_from_json reads hyperedges from the
top-level and nested slots; a full validation wipeout is reported loudly.
#2490 (thanks @PapiScholz): the skill Step-5 flow passes curated
community_labels to to_json, so graph.json ships community_name.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-05 22:08:29 +01:00
safishamsi 4e7e6b1f7e fix(extract,watch): C# assembly-aware partial merge, incremental edge preservation, extract crash recovery; bump to 0.9.33
#2411 (thanks @JensD-git): key the C# partial-class merge on assembly
(nearest ancestor .csproj/.fsproj/.vbproj) in addition to namespace and
name, so same-name partial classes in different assemblies stay distinct
while genuine partial halves within one project still merge. Fixes a
0.9.32 regression from #2332.

#2437/#2438 (thanks @aryanbonigala, builds on PR #2439): incremental
rebuilds no longer drop member-call and indirect_call edges from a
changed file into an unchanged target. Re-resolution now sees the
unchanged corpus (nodes, contains/method edges, and _callable markers,
which now persist to graph.json like _origin); edges to a genuinely
removed target are still evicted.

#2444/#2445 (thanks @Baziar, builds on PRs #2461/#2458): a
BrokenProcessPool triggers the sequential fallback instead of being
swallowed per future, a failed worker file is retried sequentially
rather than merged as empty, and a whole-pass AST failure on a fresh
build exits non-zero instead of writing a zero-node graph
(--allow-partial opts into a best-effort partial graph).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-05 01:41:12 +01:00
safishamsi 50be9dcd36 fix(cli): recover edge direction from _src/_tgt in path and explain (#2309)
graphify path decided hop direction from the persisted source/target order,
so a link stored in flipped order (pre-#563 graphs, raw dumps, merge-driver
output) printed backwards; explain had the same defect, and the query/merge
load shims clobbered in-file _src/_tgt markers. All now honor _src/_tgt (the
build-side direction truth), falling back to arc order for markerless files.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-30 18:38:58 +01:00
0bLoM 2ca565ac1a fix: explain no longer resolves an ambiguous name to an arbitrary file
`_find_node` ranks matches but never reports that a tie was broken, so callers
taking `matches[0]` present one arbitrary file as the answer. Two workspaces
that each define `MetricsPort` put both nodes in the same `exact` tier,
separated only by `G.nodes()` iteration order — reorder the graph and the same
query answers with a different file, equally confidently.

`affected` already declines this via `resolve_seed()` returning None, and
`path` warns when the top two scores are within 10%. Only `explain` (and the
MCP `get_neighbors` tool, which shares the matcher) was silent.

Split `_find_node` into `_find_node_tiers` (logic unchanged, tiers exposed) plus
a flattening wrapper, and add `find_node_ambiguity()`, which returns rivals only
when the winning tier spans multiple source files. Matches within one file (a
file node plus its members) stay ordinary precedence and resolve as before.

`_disambiguate_file_node_labels` (#2032) already relabels colliding *file*
nodes; this covers the symbol case it does not reach.
2026-07-30 17:13:48 +01:00
himanshupatro-334 6d6b674b20 fix: preserve edge direction in merge-graphs 2026-07-29 17:02:00 +01:00
safishamsi 2f78439ffd fix: load raw edges-keyed graphs + stop pruning alive files as deleted (#2212, #2210)
#2212: benchmark, the merge-driver graph load, and callflow_html crashed
or silently failed on a --no-cluster graph.json (edges stored under
'edges', not 'links'). A shared load_node_link_graph helper normalizes
links/edges before node_link_graph and is used at all three sites.

#2210: _stale_graph_sources compared graph source_file spellings to the
scan with a raw string test (no NFC), and pruned any non-match with no
liveness check, so alive files (macOS NFD paths, legacy basenames) were
pruned as 'deleted'. It now compares NFC-on-both-sides and is fail-closed:
a corpus-missing source whose file still exists is pruned only when the
exclusion is provable, else kept with a warning. Prune message corrected
to 'deleted or excluded'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 10:46:47 +01:00
safishamsi 137dcf23fe fix(extract): incremental --no-cluster merges instead of overwriting the graph (#2169)
An incremental `extract --no-cluster` wrote only the changed files over
graph.json with no merge, dropping every node/edge owned by an unchanged
file; and the id-canonicalization pass only learned batch files, so the
changed file's cross-file edges kept absolute-path target ids and
dangled. The raw path now merges the existing graph forward with the same
replace/prune semantics as the clustered path (new merge_raw_extraction
helper in build.py, shared loader), refuses to overwrite a corrupt
existing graph, and the remap pass now also learns in-root edge
target_file paths (existence-gated) so cross-file targets canonicalize.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-25 22:53:43 +01:00
safishamsi 36b5e770eb fix(detect): stop the sensitive-file filter from dropping topic docs and real source (#2106)
The heuristic over-matched and silently dropped legitimate files:
- prose `.md`/`.rst` whose topic slug ends in a keyword (privacy-tokens.md,
  token-economics.md) — only code was exempt, not prose;
- the unbounded Stage-2 `service.account` substring (regex `.` wildcard) matched
  real source (google/oauth2/service_account.py) and prose slugs.
It also MISSED real secrets (.npmrc, .pypirc, secring, .git-credentials, and
case variants on case-insensitive filesystems), which were being indexed.

Fix: move service_account/aws_credentials to the boundary-checked keyword path
(so real source is spared, downloaded key files still drop), add a prose-note
carve-out (multi-word slugs indexed, bare `secrets.md`/`token.md` still dropped),
tighten id_rsa with a left boundary, add the missed secret dotfiles + secring,
lowercase the dir/segment comparisons, and count multi-dot slugs as multi-word.
Net effect is stricter on real secrets and stops the false-positive data loss.

Traceability: `graphify extract` now names the files skipped as sensitive (not
just a count), so a wrongly-flagged file is visible.
2026-07-22 18:40:11 +01:00
safishamsi 4c759b6dd6 fix(cli): deterministic tie-break for grouped explain tail (follow-up to #2094) 2026-07-22 15:31:56 +01:00
Yyunozor 0ef452c570 fix: group cut connections by file on explain for high-degree nodes
`graphify explain "<node>"` sorts a node's connections by neighbor
degree and shows only the top 20, appending a bare "... and N more"
for the rest. On a high-degree node (a logging/error helper called
from dozens of places is typical) that leaves the exact question
explain is meant to answer - "who calls this, what's the impact?" -
unanswered for 97% of the callers, with nothing pointing at where
they live (#2009).

The top-20 list and its ordering are unchanged (no behavior change
for nodes at or under the cutoff). Past it, the cut connections are
now grouped by (direction, source_file) with counts and printed
under a "Grouped by file:" section, sorted by count descending, so
the caller/callee distribution is visible without falling back to a
repo-wide grep. The aggregation itself is capped at 20 files with
its own "... and N more files" line for the pathological case where
the remainder is spread across more files than that.

Full per-caller detail behind a flag (--callers --all / --group-by=dir
as proposed in the issue) is left for a follow-up: it's a larger,
more opinionated surface (flag naming, pagination semantics) than
the literal complaint needs, and the default output no longer hides
the answer either way.

Four regression tests in tests/test_explain_cli.py: the pre-existing
truncation notice on a 30-connection node is unchanged, the grouped-
by-file output carries real counts (3 files, one at 4 and two at 3)
that sum back to exactly the cut total (no silent loss), a
byte-for-byte no-op check for nodes at/under the cutoff, and a
boundary check pinning the cutoff itself at exactly 20 connections
(no section) vs exactly 21 (one grouped entry) — the earlier tests
sit well clear of the edge and wouldn't catch a future off-by-one.
2026-07-22 15:27:07 +01:00
Yyunozor 4ace95182b fix: preserve calls edge direction in graphify query output
`graphify query` builds an undirected nx.Graph (so BFS/DFS can explore
both callers and callees of the seed node), but its text renderer
assumed the BFS/DFS visit order (u, v) was always the edge's
(source, target). On an undirected graph that assumption only holds
when the seed happens to be the caller: seeding on the callee makes
BFS/DFS visit the callee first, so a `caller --calls--> callee` edge
was rendered backwards as `callee --calls--> caller`. graph.json's
own source/target fields stay correct on disk; only the query
rendering was wrong.

`graphify path` and `graphify explain` don't have this problem
because they force directed=True on load (#849, #853), and the MCP
query_graph tool's _load_graph() does the same. Doing that for CLI
`query` too was tried and reverted: forcing a DiGraph makes
G.neighbors() return successors only, so a query seeded on a
leaf/sink node (no outgoing edges) found zero neighbors instead of
its callers — a recall regression, not just a display fix, and it
would make the CLI and MCP query tools diverge in what they discover
even though they'd render direction identically.

Fix instead mirrors the _src/_tgt pattern graphify/build.py already
uses for the same underlying problem (undirected storage loses
direction): the CLI now stashes each link's true source/target on
its edge data as _src/_tgt before constructing the (still undirected)
graph, and _subgraph_to_text renders EDGE lines from _src/_tgt when
present, falling back to (u, v) otherwise. Traversal itself is
unchanged, so recall is unaffected — verified against the unpatched
CLI, the node counts returned for the same seeds are identical before
and after this fix, only the printed edge direction changes.

Adds two regression tests in tests/test_query_cli.py seeding the same
`calls` edge from both endpoints; the callee-seeded case fails on the
prior code with the exact backwards-edge symptom above.
2026-07-21 21:45:53 +01:00
HerenderKumar bb8f5bd675 docs(cli): surface --code-only in the extract usage and README (#2071)
The top-level `graphify --help` has listed --code-only since #1734, but the
`graphify extract` usage string and the README never mentioned it. A user
evaluating graphify on a "can this run with no network call" constraint sees
the extract usage / README first and can conclude the flag doesn't exist.

Add --code-only to the extract usage line, name it in the Privacy section and
the command reference, and add a test asserting the usage advertises it.
2026-07-21 21:33:12 +01:00
safishamsi 1fbc623271 fix(query): report real call-site lines + stop silent query truncation (benchmark)
Two correctness defects found in a head-to-head benchmark of 0.9.22.

Caller line numbers: explain/affected/get_neighbors/query printed the caller
node's def line for an incoming call, presented as a precise citation, so
click-through landed in the wrong place. The `calls` edge already carries the
true call-site line (engine.py sets it); every caller/relation listing now reads
the traversed edge's source_file:source_location, falling back to the node's own
line only when the edge lacks one.

Silent query truncation: rendered nodes were degree-ordered (a low-degree
definition node ranked last, cut first), the queried symbol wasn't guaranteed to
appear, and the truncation marker sat only at the end so silence read as absence.
Nodes are now ranked by hop distance from the seeds (deterministic), the seed the
question named is rendered first and never truncated, and a prominent TRUNCATED
notice at the top states shown/total counts and how to widen the budget. Also
rewires the seed-first ordering the renderer already supported — a branch merge
had silently dropped the `seeds=` argument, leaving it dead code.
2026-07-21 18:39:05 +01:00
safishamsi 87c870495f fix(label): --no-label placeholders no longer permanently suppress real labels (#2073)
A `cluster-only --no-label` run wrote "Community N" placeholders into
.graphify_labels.json plus a matching .sig, and the reuse path treated them as
fresh, so real labels were never regenerated on later runs. Two fixes: (a) don't
persist the labels sidecar (or its .sig) on a placeholder-only run, so a later
run generates real labels; (b) treat a stored "Community {cid}" as absent in the
reuse path so an already-polluted sidecar self-heals via the hub labeler while
genuine labels are still reused with no LLM call. The watch/update rebuild had
the same placeholder-perpetuation twin — fixed alongside.
2026-07-21 12:59:10 +01:00
safishamsi b194301c05 fix(path): deterministic route + honest edge relation, not fabricated calls (#2074)
`graphify path` (and the MCP shortest_path tool) ran shortest_path over
G.to_undirected(as_view=True), whose neighbor iteration is a hash-seeded set
union, so among equal-length paths BFS returned a route that varied per process.
Build a sorted, materialized undirected graph so the chosen path is canonical.

The hop label also printed a relation read from an arbitrarily-collapsed parallel
edge, so it could show `calls` on a pair that only carries `references`. Force
multigraph on the cli path reload so parallel links survive, and render the
ACTUAL stored relation(s) via edge_datas, falling back to an honest "related"
when the edge has none. Serve's shared graph is left untouched (its degree feeds
query-seed tie-breaks); the fix is applied locally in both path readers.
2026-07-21 12:54:58 +01:00
safishamsi 5b05f0d9ab fix(cli): disambiguate file labels on the --no-cluster extract path too (#2032)
The build_from_json label pass only covered the clustered/update paths; the raw
`extract --no-cluster` path writes the merged node list directly, so colliding
basenames stayed un-disambiguated there. Factor the logic into a shared
_file_label_reassignments core with a list-based variant
(disambiguate_file_labels_in_nodes) and apply it on the raw merged nodes.
Caught by the clean-venv edge-case battery.
2026-07-20 16:35:14 +01:00
safishamsi 1f4e3b2fe1 fix(cli): wire god-nodes subcommand + accept --output alias on extract (#2004)
Part 2: `god_nodes` was an analyzer, an MCP tool, and a README-advertised
capability, but `graphify god_nodes` errored with "unknown command". Add a
read-only `god-nodes`/`god_nodes` subcommand mirroring `affected` (--graph,
--top, --json), routing labels through sanitize_label.

Part 3: `--output DIR` on `extract` was silently dropped (output fell back to
the default dir). It is now an alias of `--out` (both space and =forms), matching
what `graphify tree` already documents. Help/usage text updated.

Part 1 (affected/reverse-dep import-id mismatch) is deferred — a build-time
id-resolution change, tracked separately.
2026-07-20 15:35:40 +01:00
safishamsi 6a56763051 fix(build): form-insensitive prune + scan-root marker (#2012)
build_merge now prunes a deleted file's nodes, edges, and hyperedges regardless
of whether their stored source_file is absolute or relative. When the caller
passed no root (the --update runbook), a node that kept an absolute path slipped
past the relative prune set and the deleted file's graph survived silently.
Matching is now form-insensitive (raw, normalized-relative, then an absolute-
identity fallback); a re-extracted file is still never pruned (#1796 preserved).

`graphify extract` also writes the .graphify_root marker after every graph write
so a later build_merge relativizes deleted-file paths correctly even under a
custom --out (its grandparent-of-graph.json fallback pointed at the wrong dir).

Regression tests in tests/test_build_merge_hyperedges_and_prune.py.
2026-07-20 13:29:54 +01:00
oleksii-tumanov c57d5711ca fix(extract): honor persisted excludes on re-extract 2026-07-20 11:31:36 +01:00
oleksii-tumanov b35898aafc fix(extract): initialize detection for pathless Postgres 2026-07-20 11:31:36 +01:00
shazeb 08166306ba fix(cache): anchor semantic cache writes to cache_root so --out round-trips (#1990, #1991)
With `graphify extract --out <dir>`, the semantic cache write and read
sides disagreed on both location and key anchoring, breaking the cache
round-trip in two ways:

- Checkpoints (#1990): `_checkpoint_chunk` called `save_semantic_cache`
  with only `root=target`, so per-chunk recovery checkpoints were written
  under `<corpus>/graphify-out/` while the reader consulted
  `<out>/graphify-out/` — creating an unwanted graphify-out/ inside the
  analyzed source tree and making every interrupted run re-extract (and
  re-bill) completed chunks.

- Final save (#1991): cli.py passed `root=out_root`, so corpus-relative
  `source_file` paths resolved against the --out directory, failed
  `p.is_file()`, and every result group was silently skipped — the cache
  the reader would consult was never populated at all, with no warning.

Fix, following the split the AST cache already uses (#1774):

- `save_semantic_cache` and `check_semantic_cache` gain a `cache_root`
  parameter mirroring `load_cached`/`save_cached`: `root` stays the
  source-key anchor (content-hash keys, source_file resolution and
  relativization), `cache_root` selects where cache files live. Omitting
  it keeps `root` for both, so existing callers are unchanged.
- `extract_corpus_parallel` plumbs `cache_root` into `_checkpoint_chunk`.
- cli.py extract passes `root=target, cache_root=out_root` at the cache
  read, the checkpoint path, and the final save, and re-anchors the
  prune sweep's live hashes to `target` (keys anchored to out_root would
  mismatch every entry and sweep the fresh cache as orphaned).
- `save_semantic_cache` now warns loudly when every result group is
  dropped because its source_file does not resolve to a real file — the
  silent-0-writes failure mode #1991 asked to surface.

Regression tests cover: checkpoint written under cache_root (not the
corpus, no corpus graphify-out/ created), recovery read finds the
checkpoint via the same root/cache_root split, the final-save call shape
writes entries where the reader looks, the all-groups-dropped warning,
and backward compatibility when cache_root is omitted.

Fixes #1990
Fixes #1991
2026-07-18 19:09:53 +01:00
shazeb 0224bcaea4 fix(install): fire the search nudge on Claude Code's Grep tool, not just Bash (#1986)
ee1df22 narrowed the Claude Code search-guard matcher from "Glob|Grep" to
"Bash" on the premise that dedicated search tools were removed and searches
go through Bash. Current Claude Code routes content search through its
first-class Grep tool (its Bash tool description actively steers away from
shell grep), so the graphify-first nudge never fired on the agent's primary
exploration path and the graph was silently bypassed.

Three-part fix, per the issue's analysis:

- Matcher: "Bash" -> "Bash|Grep" in _claude_pretooluse_hooks. Glob already
  fires the read nudge via "Read|Glob", so Grep was the only orphaned tool.
- Guard body: the hook-guard search branch only inspected tool_input.command,
  which a Grep call doesn't carry (it has pattern/path/glob). A Grep-shaped
  input (pattern present, no command) is now treated as a search — it IS one
  by definition — and nudges whenever a fresh graph exists. The Bash
  token-matching path is unchanged, and a command-carrying input never
  triggers the Grep shape, so non-search Bash calls stay silent.
- Idempotency: "Bash|Grep" added to the four install/uninstall dedup filters
  (claude + codebuddy), so upgrading replaces the stale "Bash" hook in place
  instead of appending a duplicate — verified against a pre-fix settings.json.

Tests: new regression tests feed Grep-shaped tool_input through
hook-guard search and assert the nudge (with graph), silence (without),
valid PreToolUse JSON, and no blocking; plus a guard that a non-search Bash
command with a stray pattern key does not nudge. Existing matcher assertions
updated across test_search_hook/test_install/test_claude_md/test_codebuddy/
test_hook_strict. Hook+install suites: 397 passed. Full suite: 3224 passed;
the 13 failures are pre-existing on clean v8 in this environment.

Fixes #1986
2026-07-18 19:07:16 +01:00
safishamsi caa6f9edb9 fix(extract): persist --no-gitignore instead of clobbering it (follow-up to #1979)
The PR always wrote gitignore=not no_gitignore into .graphify_build.json, so a
flag-less `graphify extract` after `--no-gitignore` reset it to True and the
git-ignored code silently disappeared again — the exact #1971 complaint. Write
False only when the flag is set (None = leave as-is, mirroring #1886 excludes),
and honor the persisted value for the run when the flag is absent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 12:09:08 +01:00
mzt006 f17e2c5638 fix: add --no-gitignore extraction opt-out 2026-07-18 12:05:00 +01:00
safishamsi 689dd6ccfd feat(hook): opt-in strict PreToolUse guard + stop crying wolf (#1840)
Agents routinely ignore the advisory "run graphify query first" nudge and read
raw files anyway. `graphify install --project --strict` (or `graphify claude
install --strict`) now installs a hook that BLOCKS the first raw source read of a
session via permissionDecision:"deny" with a redirect to graphify query, then
downgrades to the soft nudge — it fires at most once per session (atomic
per-session marker) so it can never strand the agent, and a recent
query/explain/path refreshes a stamp that suppresses it. Claude Code only;
Bash-grep and Glob stay nudge-only; Gemini/Codex/OpenCode are unchanged.
GRAPHIFY_HOOK_STRICT=1/0 toggles at runtime without a reinstall; default installs
are byte-identical (soft nudge).

Also fixes #1840 for the default soft nudge: the guard no longer fires for reads
of out-of-project files, and softens to a non-mandatory nudge when the graph is
stale for the target file. Gating is ~3 stat calls (no corpus walk) and fails
open. Begins the 0.9.19 cycle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 17:10:41 +01:00
SinghAman21 be80ee82d5 fix(extract): anchor source_file on the scan root, not the --out dir (#1941)
`graphify extract <root> --out <dir>` reduced every node's `source_file` to a
bare filename, so a graph.json could no longer be resolved back to files on
disk by joining source_file onto the scan root.

--out passes the output dir as cache_root to relocate the cache, but that value
also anchored relativization. Every scanned file then failed relative_to(root),
fell through to the #1899 out-of-root fallback, tripped its `updepth > 3`
walk-up guard -- written for a stray ProjectReference, not a whole corpus -- and
collapsed to a basename. On Windows an --out on another drive hit the
cross-drive branch and basenamed unconditionally, which is what the reporter
saw: 0 of ~120k source_files kept a separator. The directory survived only in
the node id slug, lossily (`.`, `/`, `\`, `-`, spaces all map to `_`), and no
other field carried it -- origin_file is stripped (#1516) and the export has no
file table -- so 0% of nodes resolved.

extract() now takes an explicit `root` anchor for source_file/ids/symbol
resolution, which the CLI pins to the scan root independent of where the cache
lives. This completes the cache/anchor decoupling #1774 started and matches
build(root=target), which already anchored on the scan root -- extract was the
lone component keying off --out.

cache_root keeps its fallback-anchor role, so callers that pass the scan root as
cache_root (watch, the no---out CLI path, tests) are unchanged, as are cache
location and out-of-root portability (#1899).
2026-07-17 11:35:11 +01:00
Alpha Nury 824cac7086 fix(detect): clear stale semantic_hash for dispatched-but-omitted files (#1948)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 11:30:53 +01:00
oleksii-tumanov 5a480a83a7 fix(merge-chunks): fail when no chunk validates 2026-07-17 10:54:07 +01:00
safishamsi 479f1af455 fix(cache): mark files partial on empty-parse truncation (follow-up to #1950)
The _partial item-marker approach couldn't fire on the most common truncation
shape: a mid-JSON cut parses to zero items, so a sliced document whose second
slice truncated empty was still stamped complete (the PR's own headline case).

- The adaptive-retry give-up sites now record the chunk's own source files in a
  result-level _partial_files list, independent of parsed items; it propagates
  through _merge_two / the recursion merges / _merge_into so it reaches both the
  per-chunk checkpoint and the run-level manifest stamp.
- _partial_source_files unions _partial_files with the item markers.
- save_semantic_cache seeds an empty group for a named partial file with no
  items so its entry is stamped partial, and carries a partial prev entry's flag
  forward so a later clean slice merging over it can't re-promote it to complete.
- the CLI final save now passes partial_source_files (computed before the save)
  so an empty-parse file isn't written back as a complete entry.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 00:02:15 +01:00
tpateeq 90dcb6a2ef fix(cache): don't promote truncated LLM chunks to the semantic cache as complete
A chunk whose LLM response is truncated (`finish_reason="length"`) and can't be
recovered by splitting, or that hits the adaptive-retry depth cap, returns a
partial node set. Today that set is checkpointed and written to the content-hash
semantic cache + manifest-stamped as complete, so the incomplete nodes are
served forever until the file content changes or `--force`.

Truncated give-up results are now tagged with an internal `_partial` marker.
`save_semantic_cache` stamps the affected file's entry `partial: True` (detected
from the marker or an explicit `partial_source_files` arg), and `load_cached`
treats a partial entry as a cache MISS, so the file re-dispatches and retries.
The file is also left unstamped in the manifest (like a failed chunk, #933) so
detect_incremental re-queues it on the next incremental run — not only on a full
/ `--force` / content-change run. A file sliced across chunks accumulates via a
partial-aware `merge_existing` peek (`load_cached(allow_partial=True)`) so a
truncated slice is never dropped or silently promoted to complete. Self-heals: a
later complete extraction overwrites the same key with a non-partial entry. The
marker is stripped after the final save so it never leaks into graph.json.
2026-07-16 23:48:07 +01:00
safishamsi 16fe8f3020 fix(io): route the remaining graph/manifest writers through the atomic helper (follow-up to #1952)
The atomic-write PR left several writers on the old truncate-then-write path.
Route them through write_json_atomic so a crash mid-write can't corrupt them:
- the --no-cluster raw graph.json dump (a core graph.json writer)
- merge-graphs / merge-chunks / merge-semantic output
- .graphify_analysis.json and .graphify_labels.json sidecars
- global_graph.py's global-graph.json and global-manifest.json

write_json_atomic gains an ensure_ascii flag so the raw-UTF-8 writers
(labels, merge outputs) keep byte-for-byte output. Adds tests for the
Windows PermissionError copy fallback and ensure_ascii=False.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 23:48:07 +01:00
tpateeq f38e98012d fix(io): write graph.json and manifest.json atomically
graph.json (the clustered `to_json` write and the `--no-cluster`/merge raw
dumps) and manifest.json were written with a direct `open()`/`write_text`, so a
crash, kill, or disk-full mid-write left a truncated, unparseable file that the
next load or `detect_incremental` then failed on.

Add `write_text_atomic`/`write_json_atomic` in graphify.paths (temp file in the
same directory + `os.replace`; JSON is streamed into the temp, not materialized
as one string) and route the graph.json writers (export.to_json,
cli._prune_graph_json_sources, the merge driver) plus detect.save_manifest
through them. The helper preserves the destination's mode (an atomic replace
never tightens 0644 to mkstemp's 0600), writes through a symlinked destination
(shared-output setups), and falls back to copy-then-delete on a Windows
os.replace lock — matching graphify.cache's existing atomic writer. On failure
the previous file is left intact and the temp removed. Not a power-loss
durability guarantee (no fsync, consistent with the rest of the codebase).
2026-07-16 23:42:47 +01:00
safishamsi 7d53a02704 fix(extract): fail closed on malformed existing graph + honor walk_errors (follow-up to #1951)
Two gaps the review found in the incomplete-build shrink guard:
- existing_graph_node_count() returned None ("proceed") on a present-but-
  unparseable graph.json, so the --no-cluster path could clobber a complete
  graph whose file was corrupt/mid-write. It now returns a MALFORMED_GRAPH
  sentinel and the caller fails closed, matching to_json's #479 handling.
- A walk that couldn't fully enumerate the corpus (permission-denied subtree,
  I/O error) is now treated as an incomplete extraction: detect()/
  detect_incremental() already record walk_errors; the extract path consumes
  them so a walk-truncated graph can't force-overwrite a complete one.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 23:42:47 +01:00
tpateeq 0f936850ba fix(extract): extend the incomplete-build shrink guard to the --no-cluster path
The clustered write is guarded by to_json's #479 shrink check, but the
`--no-cluster` raw-dump path writes graph.json directly and had no guard, so an
incomplete `--no-cluster` build could still overwrite a larger complete graph
with a partial one — the residual gap noted in the original change.

Add `existing_graph_node_count` in graphify.export (mirrors to_json's guard and
respects the graph-size cap) and, on the raw path, refuse the write with exit 1
before the manifest when the build was incomplete and the new graph has fewer
nodes than the existing one — unless --allow-partial is passed. Both write paths
now enforce the same guarantee.
2026-07-16 23:38:09 +01:00
tpateeq fcc01dbcdf fix(extract): don't force-write a partial graph over a complete one
A full `graphify extract` writes the final graph with `to_json(..., force=True)`,
which bypasses the #479 shrink guard. That is correct for a clean build that
legitimately shrinks (dedup collapse, deleted code), but when this run's
extraction was incomplete — an AST pass crashed, or some semantic chunks failed —
forcing the write lets a partial graph silently overwrite a good complete one.

The build now tracks incompleteness (AST-pass failure, semantic-pass crash, or
succeeded < total chunks) and falls back to the shrink guard (force=False) on an
incomplete run, so a smaller partial graph is refused rather than written. It
exits non-zero before the manifest is written, so the manifest is never stamped
for a graph we declined to write and the next run re-attempts. `--allow-partial`
restores force=True to override intentionally.

Note: the `--no-cluster` raw-dump path writes graph.json directly and has no
shrink guard; this change covers the normal clustered build path only.
2026-07-16 23:38:09 +01:00
safishamsi 3305ef1059 fix(merge-chunks): coerce non-numeric chunk token counts (follow-up to #1953)
An untrusted chunk with a non-numeric input_tokens/output_tokens would abort
the whole merge with a TypeError after other chunks had already merged. Coerce
to 0 so a bad token field can't defeat the per-chunk validation guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 23:38:09 +01:00
tpateeq ee74d80b45 fix(merge-chunks): validate untrusted subagent chunk JSON before merging
`graphify merge-chunks` concatenates agent-written `.graphify_chunk_*.json`
files with only a JSON-decode guard, so an oversized payload or a crafted
node/edge id (e.g. `../../etc/passwd`) flowed straight into the merged graph.

Route each chunk through `load_validated_semantic_fragment`, which stats the
file size BEFORE reading it (a multi-GB chunk can't blow up memory), parses the
JSON, and validates the byte/count caps + the node/edge id charset that blocks
path traversal (#825). An invalid chunk is skipped with a warning (filter
semantics; never abort). Left OUT of build_from_json/load_graph_json on purpose:
those must keep loading valid pre-existing graphs.

Also relax two over-strict checks in the shared validator that would otherwise
silently drop whole legitimate chunks (a relaxation — never a new rejection):
- file_type is no longer gated: build coerces any value via _FILE_TYPE_SYNONYMS
  (unknown -> "concept", #840), so synonyms like "markdown"/"tool" the loader
  maps must not fail validation.
- the id charset now allows Unicode word chars (build's normalize_id preserves
  CJK/Cyrillic/accented-Latin ids); the explicit path-separator/".." check still
  blocks directory escape.
Corrects the stale module docstring (the validator serves the devin skill path
and merge-chunks, not skill-opencode/codex).
2026-07-16 23:36:49 +01:00
SinghAman21 0d018b4c6d fix(cache): key semantic cache on the extraction prompt (#1939)
The semantic cache keyed entries on sha256(file content + path) alone, with
no component for the extraction prompt that produced them. After an upgrade
that changed the prompt, every unchanged file was a cache hit and replayed
the older prompt's extraction: the run exited 0, cost.json looked cheap, and
the graph silently carried two prompt generations side by side. The reporter
saw 506 of 512 docs replay an older vintage on a rebuild they expected to be
cold; deleting the whole cache was the only workaround.

output, and invalidating them on every release would re-bill extraction for
unchanged files. Fingerprinting the prompt itself keeps both properties:
entries survive releases that don't touch the prompt, and invalidate only
when it actually changed.

Semantic entries now live under cache/semantic/p{fingerprint}/, mirroring the
AST cache's v{version}/ layout. Both extraction paths pass their prompt: the
Python/CLI path from llm.py's _EXTRACTION_SYSTEM (shared by every backend),
and the skill path via a new prompt_file argument in Step B0/B3 naming the
references/extraction-spec.md the subagents were handed. The fingerprint
normalizes line endings so a CRLF checkout isn't mistaken for a new prompt.

Pre-existing entries predate fingerprinting and have unknowable vintage, so
they are still served rather than re-billing a whole corpus on upgrade — but
check_semantic_cache now warns with the count, turning "no signal at all"
into a visible one. merge_existing refuses to fuse such an entry into a
current-vintage write, which would mix two prompts inside one entry and then
attest the result to a prompt that produced half of it.

Old-fingerprint entries are pruned by liveness only, never swept wholesale
the way stale AST versions are: two hosts with different prompts (verbose vs
compact extraction-spec) can share one graphify-out/, and a wholesale sweep
would have each run delete the other's entries and re-bill on every
alternation. prune/clear/cached_files glob recursively so fingerprinted
entries can't become unprunable orphans (the #1527 failure mode).

The two monolith skills (aider, devin) inline their prompt instead of
shipping a spec sidecar and stay on the unfingerprinted path for now.
2026-07-16 16:38:12 +01:00