RepoWiki

From the RepoWiki wiki

RepoWiki is a system that generates Wikipedia-style wiki articles documenting a git repository's architecture, features and contributors using large language models. It is for software teams and maintainers who want to understand a codebase's structure and history without reading source code directly, and it solves the problem of keeping documentation synchronized with code changes. RepoWiki analyzes the repository's structure, clusters related files, generates feature pages and contributor biographies, detects when code changes make claims stale, and updates the wiki to stay current.

flowchart LR
  n1[["Client utilities"]]
  n2[["Command-line tools"]]
  n3[["Continuous integration"]]
  n4[["Core data model"]]
  n5[["Evaluation framework"]]
  n6[["File clustering"]]
  n7[["GitHub integration"]]
  n8[["In-flight tracking"]]
  n9[["Indexing CLI"]]
  n10[["Issue tracking"]]
  n11[["Link verification"]]
  n12[["LLM integration"]]
  n13[["Manifest generation"]]
  n14[["People management"]]
  n15[["Project configuration"]]
  n16[["Query interface"]]
  n17[["Repository indexing"]]
  n18[["Repository storage"]]
  n19[["Site articles"]]
  n20[["Site rendering"]]
  n21[["Wiki freshness"]]
  n22[["Wiki writing"]]
  n22 -->|"45 calls, 56 imports"| n4
  n18 -->|"47 calls, 23 imports"| n4
  n8 -->|"26 calls, 36 imports"| n4
  n21 -->|"25 calls, 28 imports"| n4
  n19 -->|"27 calls, 21 imports"| n20
  n2 -->|"21 calls, 26 imports"| n4
  n8 -->|"30 calls, 17 imports"| n2
  n14 -->|"10 calls, 33 imports"| n4
  n20 -->|"28 calls, 15 imports"| n19
  n5 -->|"20 calls, 21 imports"| n2
  n19 -->|"18 calls, 20 imports"| n4
  n20 -->|"14 calls, 19 imports"| n4
  n11 -->|"15 calls, 15 imports"| n4
  n5 -->|"9 calls, 20 imports"| n16
  n8 -->|"21 calls, 8 imports"| n19
  n14 -->|"7 calls, 21 imports"| n17
  n16 -->|"12 calls, 10 imports"| n4
  n8 -->|"14 calls, 7 imports"| n20
  n22 -->|"21 imports"| n17
  n2 -->|"18 imports"| n12
  n13 -->|"3 calls, 12 imports"| n4
  n14 -->|"7 calls, 7 imports"| n2
  n22 -->|"14 imports"| n12
  n5 -->|"2 calls, 11 imports"| n4
  n21 -->|"13 imports"| n17
  n2 -->|"6 calls, 6 imports"| n8
  n5 -->|"12 imports"| n12
  n8 -->|"11 imports"| n12
  n8 -->|"10 imports"| n17
  n14 -->|"9 imports"| n18
  n12 -->|"1 call, 7 imports"| n4
  n21 -->|"8 imports"| n6
  n21 -->|"8 imports"| n22
  n22 -->|"8 imports"| n13
  n21 -->|"7 imports"| n12
  n8 -->|"6 imports"| n21
  n8 -->|"6 imports"| n22
  n13 -->|"6 imports"| n6
  n13 -->|"6 imports"| n12
  n22 -->|"6 imports"| n11
  n5 -->|"3 calls, 2 imports"| n17
  n8 -->|"2 calls, 3 imports"| n16
  n14 -->|"5 imports"| n22
  n21 -->|"5 imports"| n18
  n1 -->|"2 calls, 2 imports"| n20
  n5 -->|"3 calls, 1 import"| n20
  n6 -->|"4 imports"| n17
  n13 -->|"4 imports"| n17
  n20 -->|"2 calls, 2 imports"| n8
  n2 -->|"1 call, 2 imports"| n14
  n7 -->|"3 imports"| n4
  n8 -->|"3 imports"| n18
  n11 -->|"3 imports"| n12
  n14 -->|"3 imports"| n12
  n14 -->|"3 imports"| n13
  n19 -->|"1 call, 2 imports"| n8
  n21 -->|"3 imports"| n11
  n21 -->|"3 imports"| n13
  n22 -->|"3 imports"| n18
  n2 -->|"2 imports"| n17
  n2 -->|"2 imports"| n18
  n2 -->|"2 imports"| n21
  n2 -->|"2 imports"| n22
  n6 -->|"2 imports"| n4
  n7 -->|"2 imports"| n17
  n8 -->|"2 imports"| n11
  n8 -->|"2 imports"| n13
  n9 -->|"2 imports"| n2
  n12 -->|"1 call, 1 import"| n18
  n13 -->|"2 imports"| n18
  n16 -->|"1 call, 1 import"| n17
  n17 -->|"2 imports"| n4
  n17 -->|"2 imports"| n14
  n18 -->|"1 call, 1 import"| n2
  n18 -->|"2 imports"| n12
  n18 -->|"1 call, 1 import"| n14
  n2 -->|"1 import"| n6
  n2 -->|"1 import"| n7
  n2 -->|"1 import"| n11
  n2 -->|"1 import"| n13
Each box is a feature; an arrow points from the feature whose code calls or imports to the one it uses.

Purpose and features

The system indexes a repository's source code, extracting symbols, imports, calls and file relationships to build a queryable model of its structure. (see Repository indexing) It partitions source files into cohesive clusters based on imports and co-changes, then assigns each cluster to a documented feature through LLM classification. (see File clustering, Manifest generation) RepoWiki generates feature pages in Wikipedia style, with claims backed by citations to specific lines of code and organized into sections for purpose, architecture and usage. (see Wiki writing) The system writes biographical narratives of contributors from git history, grouping their work by time and pull request and tracking their ownership and activity. (see People management)

It detects when code changes would make wiki claims stale, automatically remaps citations through diffs, and rewrites pages to reflect the current repository state. (see Wiki freshness) RepoWiki serves the wiki as a static website with article search, history pages showing how documentation evolved, and a main page with featured articles and open pull request summaries. (see Site rendering, Site articles) It displays open pull requests on the wiki, showing which claims they would change and auto-generating summaries of their code impact within a token budget using an LLM. (see In-flight tracking) An evaluation framework measures how well the wiki answers questions about the repository, comparing its accuracy and efficiency against a baseline code-querying agent. (see Evaluation framework)

Request paths

  • When building a wiki from scratch, the indexing CLI calls Repository indexing to parse the repository and write a JSON index, which Command-line tools reads and passes to File clustering to partition files, then to Manifest generation to assign features via LLM.[1]
  • When writing pages, Wiki writing builds context packs from the repository and previous revisions, sends them to LLM integration for LLM calls, verifies each claim's citations against the index, and stores the finished pages in Repository storage.[2]
  • When updating to a new commit, Wiki freshness detects membership drift against stored features, uses import and co-change signals to assign new files, remaps stale citations through diffs, and calls Wiki writing to rewrite affected pages.[3]
  • When fetching in-flight pull requests, In-flight tracking reads issues and pulls via GitHub integration, diffs the pull request head against the wiki's base, sends the diff to LLM integration for impact summaries, and marks stale claims.[4]
  • When rendering the site, Site rendering exports the wiki from Repository storage, passes it to Site articles to generate HTML for pages and history, and serves the static site.[5]
  • When evaluating the wiki, Evaluation framework reads pages via Query interface's search and read tools, calls an agent through LLM integration, and grades answers against a held-out question set.[6]

Feature dependencies

  • People management depends on Repository indexing for author and ownership data and Repository storage to persist person narratives.[7]
  • Evaluation framework depends on Query interface to search and read pages, LLM integration to call agents, and Command-line tools to access repository tools.[8]
  • Wiki writing depends on Core data model for schema validation, Repository indexing for code location lookups, LLM integration for LLM calls, and Manifest generation for feature context.[9]
  • Wiki freshness depends on Repository indexing for file change analysis, File clustering for drift detection, Manifest generation for operations, and Wiki writing to rewrite pages.[10]
  • In-flight tracking depends on Repository indexing for diff and impact analysis, GitHub integration to fetch pull requests, LLM integration to summarize changes, and Wiki freshness to detect staleness.[11]
  • Site rendering and Site articles depend on Core data model for page structure and Repository storage for the wiki export.[5]
  • Manifest generation depends on File clustering to group files and LLM integration to assign features to clusters.[12]
  • Link verification depends on LLM integration to validate Wikipedia links and on Wiki writing and Wiki freshness to rewrite links.[13]
  • Command-line tools orchestrate all other features, importing and calling Repository indexing, Manifest generation, Wiki writing, Wiki freshness, People management, and Repository storage.[14]
  • In-flight tracking feeds into Site rendering to display open pull requests, merging in-flight data with the rendered wiki.[15]
  • Repository storage depends on LLM integration to persist batch requests and on Core data model for schema validation.[16]
  • GitHub integration depends on Repository indexing to resolve repository identities from git remotes and on Core data model to structure snapshot data.[17]

Infrastructure

  • Continuous integration runs type checking, linting, and tests on every pull request and push to main, enforcing code quality through GitHub Actions. (see Continuous integration)
  • Project configuration sets up TypeScript, Node, and Vitest for the monorepo across packages managed by pnpm workspaces. (see Project configuration)
  • Build tools including Biome for linting and formatting, and Astro for static site generation, are used by the build scripts and site rendering layer. (see Project configuration, Site rendering)
  • Issue templates provide structured forms for reporting accuracy issues, bugs, decisions, features, and evaluation sessions, guiding contributors through the project's quality gates. (see Continuous integration)

References

  1. ^ scripts/index-repo.ts:L20@3993d10 (count)
  2. ^ packages/engine/src/verify/claims.ts:L11@3993d10
  3. ^ packages/engine/src/freshness/drift-call.ts:L12@3993d10
  4. ^ packages/engine/src/inflight/effects.test.ts:L8@3993d10
  5. ^ a b packages/site/src/base.test.ts:L3@3993d10
  6. ^ packages/eval/src/accuracy.ts:L9@3993d10
  7. ^ packages/engine/src/people/ownership.test.ts:L1-5@3993d10
  8. ^ packages/eval/src/agent.test.ts:L1-5@3993d10
  9. ^ packages/engine/src/verify/architecture.test.ts:L1@3993d10
  10. ^ packages/engine/src/freshness/drift-call.ts:L3@3993d10
  11. ^ packages/engine/src/inflight/effects.ts:L9@3993d10
  12. ^ packages/engine/src/manifest/build.ts:L4@3993d10
  13. ^ packages/engine/src/link/wikipedia.cassette.test.ts:L2@3993d10
  14. ^ packages/engine/src/index.ts:L91@3993d10
  15. ^ packages/site/src/main-page.ts:L3@3993d10
  16. ^ packages/engine/src/store/batch-requests.test.ts:L6@3993d10
  17. ^ packages/engine/src/github/identity.ts:L2@3993d10

This page was last edited on 9 October 2026, at commit 3993d10 (PR #619).