Methodology
Calculated, not generated.
What RepoInsight reads, how it turns that into numbers, and where those numbers stop being useful.
How it works
- 01
Read
Repository metadata, default-branch commits, contributors, releases, issues and pull requests are requested from GitHub's public REST API.
- 02
Calculate
Pure functions turn that data into figures over the last 90 days, or the longest recent period that could be read in full. The same data always produces the same report.
- 03
Show the evidence
Every number carries the inputs it came from, the formula, what it suggests and what it does not prove.
Rules we hold to
- No composite score
- Signals are never rolled into one health number. There is no defensible way to weigh them for every project.
- Stated thresholds
- Every yes or no on the checklist comes from one fixed rule, and the report prints that rule.
- Deterministic
- No language model writes any part of a report. Answers come from fixed rules over calculated figures.
- Honest coverage
- Very busy repositories cannot be read in full. A partial count can prove a yes, because it is a lower bound, but never a no. Anything else is marked not enough data.
For contributors
- Where to start
- Open issues labelled good first issue or help wanted that nobody is assigned to. Finding a first task is the barrier newcomers report most.
- Will it be merged
- The share of pull requests from outside the team that were merged rather than closed, and how long merging took.
- Will anyone reply
- The median wait for a first reply from another person. Bot comments never count, because bots often reply first.
- Is the process written down
- Whether GitHub detects a README, contributing guide, code of conduct, licence and templates.
These follow published research on newcomer barriers, pull request abandonment and first-response times. Sources are listed in the project README.
Limitations
Only public data is visible. Private forks, internal trackers and chat never show up.
Commit metrics cover the default branch. Long-lived release branches are not included.
Very active repositories exceed the collection limit; the report states the period it covers.
Authorship follows GitHub attribution. Squash merges and unlinked emails blur who did the work.
Projects that publish through tags or a package registry show no GitHub Releases.
Response times see comments and merges; a pull request answered only by an approving review looks unanswered.
Tone is not measured. Nothing here tells you whether a community is welcoming.
Activity is not quality. A quiet repository can be finished; a busy one can be unstable.
Privacy
RepoInsight retrieves only what GitHub already serves publicly for the repository you enter. It stores nothing, sets no cookies and asks for no credentials. Each report also documents its own formulas under every metric.
See it in the example report →