From b49b14f6f702258aae088fefe9c354018aae55cd Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 09:30:45 +0300 Subject: [PATCH 01/10] Rename README.md to grimoirelab-metrics.md moving this content to new file to better focus on grimoirelab --- README.md => grimoirelab-metrics.md | 0 1 file changed, 0 insertions(+), 0 deletions(-) rename README.md => grimoirelab-metrics.md (100%) diff --git a/README.md b/grimoirelab-metrics.md similarity index 100% rename from README.md rename to grimoirelab-metrics.md From cb9c7e1e7acac91c2a0070c2b0da1b8983bda8ae Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 09:59:58 +0300 Subject: [PATCH 02/10] Update grimoirelab.md remove metrics definitions to be included in new METRICS.md --- grimoirelab-metrics.md | 198 ----------------------------------------- 1 file changed, 198 deletions(-) diff --git a/grimoirelab-metrics.md b/grimoirelab-metrics.md index 666b956..cef72d7 100644 --- a/grimoirelab-metrics.md +++ b/grimoirelab-metrics.md @@ -143,201 +143,3 @@ docker run --rm \ --grimoirelab-project "npm-popular-components" \ --output metrics.json ``` - -## Project Health Metrics - -This is the list of the metrics generated by this tool: - -### Pony factor - -The metric is defined as the number of individuals, who produce up to the -first 50% of the total number of code contributions (in descending order) -within a given time period. - -A low Pony Factor implies a high dependency on these individuals, making the -project vulnerable if they were to leave. - -*Also known as: Lottery Factor, Bus Factor, Contributor Absence Factor.* - -CHAOSS definition of the [Contributor Abscence Factor](https://www.chaoss.community/kb/metric-contributor-absence-factor/). - -### Elephant factor - -The metric is defined as the number of unique organizations producing up to the -first 50% of the total number of code contributions (in descending order) -within a given time period. - -It was first defined by Bitergia, and it applies the concept of the -Pony Factor metric and takes it to contributing Organizations. - -Contributions are focused on Git commits, and the organization -is determined by the email address of the commit author. - -CHAOSS definition of the [Elephant Factor](https://www.chaoss.community/kb/metric-elephant-factor/). - -### Number of contributing organizations - -This metric quantifies the total number of distinct organizations whose members -have made contributions to an open source project over a specified period. - -Contributions are focused on Git commits, and the organization -is determined by the email address of the commit author. - -CHAOSS definition of the [Organizational Diversity](https://www.chaoss.community/kb/metric-organizational-diversity/). - -### Number of organizations contributing recently - -This metric quantifies the number of unique organizations whose members have -actively made contributions to an open source project within the last 90 days. - -Contributions are focused on Git commits, and the organization -is determined by the email address of the commit author. - -CHAOSS definition of the [Organizational Diversity](https://www.chaoss.community/kb/metric-organizational-diversity/) -filtered by a specific timeframe. - -### Number of recent contributors - -This metric quantifies the total count of unique individuals who contributed -within the last 90 days. - -CHAOSS definition of [Contributors](https://www.chaoss.community/kb/metric-contributors/) filtered by a -specific timeframe. - - -### Number of recent commits - -This metric counts the total number of commits made to the project within the -last 90 days. - -CHAOSS definition of [Code Changes Commits](https://www.chaoss.community/kb/metric-code-changes-commits/) filtered by -the last 90 days. - -### Contributor Growth Rate - -This metric measures the growth rate of active contributors, defined as the -number of people sending one or more code contributions (in this case, -Git commits) in a given period. - -To calculate it, the period is split into two halves, and the number of active -contributors in each half is compared. The growth rate is the difference -between the second half and the first half, divided by the number of active -contributors in the first half. - -```math -GrowthRate (t_1, t_2) = \frac{C_a(t_2) - C_a(t_1)}{C_a(t_1)} -``` - -### Contributor Growth - -This metric measures the growth of active contributors, defined as the -number of people sending one or more code contributions (in this case, -Git commits) in a given period. - -To calculate it, the period is split into two halves, and the number of active -contributors in each half is compared. Growth is the difference between the -second half and the first half. - -```math -Growth (t_1, t_2) = C_a(t_2) - C_a(t_1) -``` - -### Number of active branches - -This metric refers to the count of branches within a project's version control -repository (for this case, Git) that have seen recent development activity, -usually indicated by new commits. - -CHAOSS definition of [Branch Lifecycle](https://www.chaoss.community/kb/metric-branch-lifecycle/). - -### Days since last commit - -This metric shows the number of days since the last commit was submitted to the -repository or the project. - -CHAOSS definition of [Code Changes Commits](https://www.chaoss.community/kb/metric-code-changes-commits/) -taking into account the last commit activity. - -### Presence of an adopters file in a standard location - -This metric verifies the existence of a file containing the full text of -the project's adopters. - -Standard practice dictates that this file is named `ADOPTERS`, `ADOPTERS.md`, -or `ADOPTERS.txt`, and is located in the root directory of the project's -source code repository. - -### Presence of a license file in a standard location - -This metric verifies the existence of a file containing the full text of -the project's chosen open source license. - -Standard practice dictates that this file is named `LICENSE`, `LICENSE.md`, -`LICENSE.txt`, or `COPYING` (a convention historically used by GNU projects) -and is located in the root directory of the project's source code repository. - -CHAOSS definition of [License Declared](https://www.chaoss.community/kb/metric-licenses-declared/). - -### Rate of contributors contributing infrequently vs. regularly - -This metric involves categorizing contributors based on the frequency, -consistency, and intensity of their contributions over a defined period. It aims -to distinguish between individuals who contribute episodically -(infrequent contributors) and those who engage with the project consistently -and often substantially (regular or core contributors). - -The CHAOSS community, for instance, defines -"[Occasional Contributors](https://www.chaoss.community/kb/metric-occasional-contributors/)" as -"people who make contributions to a project on an irregular basis". - -### Number of contributors who have contributed in previous periods - -This metric, often referred to as "returning contributors" or as an indicator -of "contributor retention," counts the number of unique individuals who were -active contributors in the last 90 days and had also made contributions in one -or more defined previous periods. - -### Rate of commits over specified periods - -This metrics calculates the rate of commits between the last 90 days and -the last year. The purpose of this metrics is having an estimation of how -distributed contributions are, and the "momentum" of the project. - -### Number of commits per repository - -Activity for each of the Git repositories analyzed. - -CHAOSS definition of [Code Changes Commits](https://www.chaoss.community/kb/metric-code-changes-commits/) filtered -by repository. - -### Number of developers per repository - -Unique participants producing commits in a given repository. - -CHAOSS definition of [Contributors](https://www.chaoss.community/kb/metric-contributors/) filtered -by repository. - -### File type metrics (code, documentation, or others) - -Type of activity done by developers, mainly split into code, documentation, and others. - -### Commit size metrics (added and removed lines) - -This provides an overview of the usual size of the code review processes and good practices -when submitting code. - -CHAOSS uses the metric [Change Request Commits](https://www.chaoss.community/kb/metric-change-request-commits/) -as a way to use this as a filter. - -### Message size metrics (total, mean, and median) - -This metrics provides information about how extensive a developer is providing information about the -change in the commit. - -### Frequency metrics for commits - -This allows to understand the consistency of developers when producing code. - -### Developer categories (core, regular and casual) - -This metric structures the developers by their activity and consistency. The core developers produce From 3e77e445ffe726037111bc2dc1a2a7b79dda0d58 Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 11:13:58 +0300 Subject: [PATCH 03/10] Create new README.md focus on explaining the project and getting started FIXME: add graphics --- README.md | 150 ++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 150 insertions(+) create mode 100644 README.md diff --git a/README.md b/README.md new file mode 100644 index 0000000..940373e --- /dev/null +++ b/README.md @@ -0,0 +1,150 @@ +# HealthyCode +Open source project health metrics and scores, built on an open model, open methodology, and open data. + +HealthyCode tells you how healthy the open source projects you depend on are, so you can make informed, data-driven decisions about which components to adopt, upgrade, or replace. It reads a project's development history (commits, contributors, and how activity changes over time) and turns it into a set of health metrics and one health score. Every score comes with the metrics behind it, so you can see why a package scored the way it did and apply your own thresholds. + +HealthyCode uses [GrimoireLab](https://github.com/chaoss/grimoirelab) to collect data, [ScanCode.io](https://github.com/aboutcode-org/scancode.io) to run pipelines, and [PurlDB](https://github.com/aboutcode-org/purldb) to store the data. + +## Why project health matters +Most software is assembled from open source packages. Security scanners are good at flagging a package with a known vulnerability, but they say nothing about a package whose last maintainer has quietly moved on... + +> HealthyCode starts with npm. Other ecosystems will follow. +> +> npm is the largest package registry in the world, and its packages depend on each other heavily. A small library can sit underneath thousands of projects. +> +> In npm, this risk spreads quickly. One unmaintained library can become a single point of failure for everything built on it, and it is also easier to take over. When a maintainer's account or email domain lapses, an attacker can claim it and publish a malicious release. + +HealthyCode helps you spot these packages before you depend on them, and keep watching the ones you already use. + +## What you get +For each package, HealthyCode returns: +- Health metrics from the project's Git history, such as recent commits, contributor activity and growth, days since last commit, pony factor (how many people do half the work), and the elephant factor (how many organizations do half the work). The full list is in [METRICS.md](METRICS.md). +- A health score between 0 and 1 that is the model's estimate of how likely the project is to be unhealthy. 0 means healthy and 1 means unhealthy. +- The settings used for the run: time window, thresholds, and the versions of HealthyCode and the scoring model. + +Results are JSON, so you can feed them into a dashboard, a report, or an allow/block policy in your own tooling. HealthyCode gives you the evidence, but the policy decision stays with you. + +Here is a trimmed result for the `semver` package: +```json +{ + "packages": { + "SPDXRef-Package-semver-954": { + "repository": "https://github.com/npm/node-semver.git", + "metrics": { + "recent_commits": 52, + "recent_contributors": 17, + "returning_contributors": 3, + "pony_factor": 2, + "elephant_factor": 1, + "days_since_last_commit": 1, + "found_file_license": 1 + }, + "score": { + "value": 0.0, + "metadata": { + "ecosystem": "npm", + "model": "health", + "version": "0.2" + } + } + } + } +} +``` + +## Who it's for +- Engineering teams deciding whether to adopt, upgrade, replace, or drop a dependency +- Open source program offices (OSPO) and supply chain teams tracking the packages their organization relies on +- Maintainers who want an outside view of their own project +- Researchers studying open source sustainability + +## When not to use it +HealthyCode measures project health. It does not scan for vulnerabilities or check license compliance. For those, see [VulnerableCode](https://github.com/aboutcode-org/vulnerablecode), [ScanCode.io](https://github.com/aboutcode-org/scancode.io), and +[ScanCode Toolkit](https://github.com/aboutcode-org/scancode-toolkit). + +## Getting started +HealthyCode needs a running [GrimoireLab 2.x](https://github.com/chaoss/grimoirelab/blob/2.x/README.md) +instance with OpenSearch. GrimoireLab collects the data, and HealthyCode +turns it into metrics and a score. + +The quickest way to run HealthyCode is with the published Docker image: + +```bash +docker run --rm ghcr.io/aboutcode-org/healthycode:0.2.0 \ + /opt/healthycode/.venv/bin/grimoirelab-metrics \ + https://github.com/npm/node-semver.git \ + --grimoirelab-url http://your-grimoirelab:8000 \ + --grimoirelab-user USER --grimoirelab-password PASSWORD \ + --grimoirelab-ecosystem npm --grimoirelab-project my-packages \ + --opensearch-url https://your-opensearch:9200 \ + --opensearch-user USER --opensearch-password PASSWORD \ + > metrics.json +``` + +Replace the URLs and credentials with your own. To analyze an SPDX SBOM instead of a single repository, mount the file into the container and pass its path. The first run for a repository takes longer because GrimoireLab has to collect its full history. + +Installation, setup, and command-line usage guidance for GrimoireLab is in [grimoirelab-metrics.md]/(grimoirelab-metrics.md). + +## How it works + +1. You give HealthyCode a Git repository URL, or an SPDX SBOM that lists Git repositories. +2. HealthyCode asks GrimoireLab to collect each repository's history. Repositories GrimoireLab hasn't seen yet are added and analyzed. +3. When the data is ready, HealthyCode computes the metrics over a time window. The default is 12 months. +4. The npm health model turns those metrics into a score. +5. The metrics, score, and run settings are written to a JSON file. + +HealthyCode runs on your own infrastructure. The only outside services it contacts are the public code hosts and registries the data comes from. + +## Access health data through ScanCode.io pipelines +You can run HealthyCode from ScanCode.io, next to your other scans: https://github.com/aboutcode-org/scancode.io/blob/main/scanpipe/pipelines/scan_repo_health.py + +The `scan_repo_health` pipeline takes a Git repository URL, collects the data with GrimoireLab, and saves the metrics and score with the project results. Your ScanCode.io instance needs access to a GrimoireLab instance to run it. + +## Access health data by Package-URL (PURL) +We are adding health data to [PurlDB](https://github.com/aboutcode-org/purldb) via API: https://health.purldb.io/api/ + +You will be able to ask for a package's health by its PURL (for example, `pkg:npm/semver`) and get the same JSON back. If the package has already been analyzed, the answer comes back right away. If not, the request starts a ScanCode.io pipeline to analyze it. Results will refresh when new versions are released. + +## The methodology +HealthyCode uses the Goal-Question-Metric (GQM) approach. You start with what you care about (the goal), then work out what you need to know (the questions), and only then choose what you can measure (the metrics). For example: + +| Goal | Question | Metrics | +| --- | --- | --- | +| The project stays maintained | Would it survive losing its top contributor? | Bus factor (CHAOSS Contributor Absence Factor)
Top author's share of commits over the last 12 months
Time since a second publisher shipped a release | + +GitHub stars are not on the list. They answer a question about popularity, and that is not what the model asks. + +# The model +The v0.2 model was trained on about 1,150 popular npm packages drawn from Census II, Census III, and deps.dev. An open source expert reviewed 200 and classified 166 as healthy or unhealthy without seeing any scores. A logistic regression then learned which metrics best separate the two groups and how much weight each one gets. The data, notebook, a workflow diagram, and the expert's classifications guidelines are in [model/npm](model/npm). + +This is a first version. Weights and thresholds will be tuned as more packages and ecosystems are analyzed, and new questions and metrics will be added, drawing on [CHAOSS](https://www.chaoss.community/), [OpenSSF Scorecard](https://github.com/ossf/scorecard), and foundation maturity models. + +## How HealthyCode compares +People have been working on open source health from three directions: +1. Organizations improving how they take in open source, through work such as OpenChain, the FINOS Open Source Maturity Model, the Good Governance Initiative, and the TODO Group. +2. Foundations grading their own projects, such as the Apache Software Foundation maturity model, the Eclipse Foundation Development Process, and the CNCF project lifecycle (sandbox, incubating, graduated, archived). +3. Community projects that work across ecosystems. CHAOSS defines open health metrics and models, and OpenSSF Scorecard scores security practices. + +Several commercial companies also sell package health scores. In most cases, you get the score but not the underlying data or the exact weights, which makes a result hard to check or reproduce. + +## HealthyCode is open code, open data, , every result can be audited +The code, the collected data, and the model weights are public. Each score comes with the metrics behind it and the settings used to produce it, and each metric comes from public commit history you can check yourself. You can also rerun the analysis on your own infrastructure to verify a result. Metrics follow CHAOSS definitions where they exist, and data collection is done by GrimoireLab, a CHAOSS project. + +HealthyCode also fits into the AboutCode stack. You can run it as a ScanCode.io pipeline and look up results in PurlDB by PURL, the standard package identifier used in SBOMs and vulnerability databases and across software supply chains. That puts a package's health data next to its license and origin data from ScanCode for more comprehensive visibility into the packages you use. + +The full state-of-the-art review is available here: + +## Project status +HealthyCode is at version 0.2.0 and under active development. npm is the first ecosystem, and the metrics and model will change as more packages are analyzed. If a result looks wrong to you, please open an issue with the package name and what you expected. That feedback goes straight into improving the model. + +## Part of AboutCode +HealthyCode is part of [AboutCode](https://aboutcode.org), a family of open source tools, open data, and open standards for healthy and secure software supply chains, alongside ScanCode, VulnerableCode, PurlDB, and Package-URL. It is developed by AboutCode and community contributors like you! + +## Contributing +Issues and pull requests are welcome. Please follow the [code of conduct](CODE_OF_CONDUCT.rst). + +## License +The code is licensed under GPL-3.0-or-later. See [LICENSE](LICENSE). + +Data produced by HealthyCode is licensed under +[CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). From c824853b49eec374e36518907cdcc9cc9d2e6843 Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 11:14:39 +0300 Subject: [PATCH 04/10] fix header typo --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 940373e..76c751a 100644 --- a/README.md +++ b/README.md @@ -114,7 +114,7 @@ HealthyCode uses the Goal-Question-Metric (GQM) approach. You start with what yo GitHub stars are not on the list. They answer a question about popularity, and that is not what the model asks. -# The model +## The model The v0.2 model was trained on about 1,150 popular npm packages drawn from Census II, Census III, and deps.dev. An open source expert reviewed 200 and classified 166 as healthy or unhealthy without seeing any scores. A logistic regression then learned which metrics best separate the two groups and how much weight each one gets. The data, notebook, a workflow diagram, and the expert's classifications guidelines are in [model/npm](model/npm). This is a first version. Weights and thresholds will be tuned as more packages and ecosystems are analyzed, and new questions and metrics will be added, drawing on [CHAOSS](https://www.chaoss.community/), [OpenSSF Scorecard](https://github.com/ossf/scorecard), and foundation maturity models. From 119463bfa95e6f47c5a4f4954999b3f464b41237 Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 11:19:18 +0300 Subject: [PATCH 05/10] Create METRICS.md separating from grimoirelab guide --- METRICS.md | 197 +++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 197 insertions(+) create mode 100644 METRICS.md diff --git a/METRICS.md b/METRICS.md new file mode 100644 index 0000000..f4f77bb --- /dev/null +++ b/METRICS.md @@ -0,0 +1,197 @@ +## Project Health Metrics + +This is the list of the metrics generated by this tool: + +### Pony factor + +The metric is defined as the number of individuals, who produce up to the +first 50% of the total number of code contributions (in descending order) +within a given time period. + +A low Pony Factor implies a high dependency on these individuals, making the +project vulnerable if they were to leave. + +*Also known as: Lottery Factor, Bus Factor, Contributor Absence Factor.* + +CHAOSS definition of the [Contributor Abscence Factor](https://www.chaoss.community/kb/metric-contributor-absence-factor/). + +### Elephant factor + +The metric is defined as the number of unique organizations producing up to the +first 50% of the total number of code contributions (in descending order) +within a given time period. + +It was first defined by Bitergia, and it applies the concept of the +Pony Factor metric and takes it to contributing Organizations. + +Contributions are focused on Git commits, and the organization +is determined by the email address of the commit author. + +CHAOSS definition of the [Elephant Factor](https://www.chaoss.community/kb/metric-elephant-factor/). + +### Number of contributing organizations + +This metric quantifies the total number of distinct organizations whose members +have made contributions to an open source project over a specified period. + +Contributions are focused on Git commits, and the organization +is determined by the email address of the commit author. + +CHAOSS definition of the [Organizational Diversity](https://www.chaoss.community/kb/metric-organizational-diversity/). + +### Number of organizations contributing recently + +This metric quantifies the number of unique organizations whose members have +actively made contributions to an open source project within the last 90 days. + +Contributions are focused on Git commits, and the organization +is determined by the email address of the commit author. + +CHAOSS definition of the [Organizational Diversity](https://www.chaoss.community/kb/metric-organizational-diversity/) +filtered by a specific timeframe. + +### Number of recent contributors + +This metric quantifies the total count of unique individuals who contributed +within the last 90 days. + +CHAOSS definition of [Contributors](https://www.chaoss.community/kb/metric-contributors/) filtered by a +specific timeframe. + + +### Number of recent commits + +This metric counts the total number of commits made to the project within the +last 90 days. + +CHAOSS definition of [Code Changes Commits](https://www.chaoss.community/kb/metric-code-changes-commits/) filtered by +the last 90 days. + +### Contributor Growth Rate + +This metric measures the growth rate of active contributors, defined as the +number of people sending one or more code contributions (in this case, +Git commits) in a given period. + +To calculate it, the period is split into two halves, and the number of active +contributors in each half is compared. The growth rate is the difference +between the second half and the first half, divided by the number of active +contributors in the first half. + +```math +GrowthRate (t_1, t_2) = \frac{C_a(t_2) - C_a(t_1)}{C_a(t_1)} +``` + +### Contributor Growth + +This metric measures the growth of active contributors, defined as the +number of people sending one or more code contributions (in this case, +Git commits) in a given period. + +To calculate it, the period is split into two halves, and the number of active +contributors in each half is compared. Growth is the difference between the +second half and the first half. + +```math +Growth (t_1, t_2) = C_a(t_2) - C_a(t_1) +``` + +### Number of active branches + +This metric refers to the count of branches within a project's version control +repository (for this case, Git) that have seen recent development activity, +usually indicated by new commits. + +CHAOSS definition of [Branch Lifecycle](https://www.chaoss.community/kb/metric-branch-lifecycle/). + +### Days since last commit + +This metric shows the number of days since the last commit was submitted to the +repository or the project. + +CHAOSS definition of [Code Changes Commits](https://www.chaoss.community/kb/metric-code-changes-commits/) +taking into account the last commit activity. + +### Presence of an adopters file in a standard location + +This metric verifies the existence of a file containing the full text of +the project's adopters. + +Standard practice dictates that this file is named `ADOPTERS`, `ADOPTERS.md`, +or `ADOPTERS.txt`, and is located in the root directory of the project's +source code repository. + +### Presence of a license file in a standard location + +This metric verifies the existence of a file containing the full text of +the project's chosen open source license. + +Standard practice dictates that this file is named `LICENSE`, `LICENSE.md`, +`LICENSE.txt`, or `COPYING` (a convention historically used by GNU projects) +and is located in the root directory of the project's source code repository. + +CHAOSS definition of [License Declared](https://www.chaoss.community/kb/metric-licenses-declared/). + +### Rate of contributors contributing infrequently vs. regularly + +This metric involves categorizing contributors based on the frequency, +consistency, and intensity of their contributions over a defined period. It aims +to distinguish between individuals who contribute episodically +(infrequent contributors) and those who engage with the project consistently +and often substantially (regular or core contributors). + +The CHAOSS community, for instance, defines +"[Occasional Contributors](https://www.chaoss.community/kb/metric-occasional-contributors/)" as +"people who make contributions to a project on an irregular basis". + +### Number of contributors who have contributed in previous periods + +This metric, often referred to as "returning contributors" or as an indicator +of "contributor retention," counts the number of unique individuals who were +active contributors in the last 90 days and had also made contributions in one +or more defined previous periods. + +### Rate of commits over specified periods + +This metrics calculates the rate of commits between the last 90 days and +the last year. The purpose of this metrics is having an estimation of how +distributed contributions are, and the "momentum" of the project. + +### Number of commits per repository + +Activity for each of the Git repositories analyzed. + +CHAOSS definition of [Code Changes Commits](https://www.chaoss.community/kb/metric-code-changes-commits/) filtered +by repository. + +### Number of developers per repository + +Unique participants producing commits in a given repository. + +CHAOSS definition of [Contributors](https://www.chaoss.community/kb/metric-contributors/) filtered +by repository. + +### File type metrics (code, documentation, or others) + +Type of activity done by developers, mainly split into code, documentation, and others. + +### Commit size metrics (added and removed lines) + +This provides an overview of the usual size of the code review processes and good practices +when submitting code. + +CHAOSS uses the metric [Change Request Commits](https://www.chaoss.community/kb/metric-change-request-commits/) +as a way to use this as a filter. + +### Message size metrics (total, mean, and median) + +This metrics provides information about how extensive a developer is providing information about the +change in the commit. + +### Frequency metrics for commits + +This allows to understand the consistency of developers when producing code. + +### Developer categories (core, regular and casual) + +This metric structures the developers by their activity and consistency. The core developers produce From 35684b1ef429b88800b3e1067bca8e1a549ce3f0 Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 11:20:04 +0300 Subject: [PATCH 06/10] Rename grimoirelab-metrics.md to grimoirelab-guide.md --- grimoirelab-metrics.md => grimoirelab-guide.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) rename grimoirelab-metrics.md => grimoirelab-guide.md (98%) diff --git a/grimoirelab-metrics.md b/grimoirelab-guide.md similarity index 98% rename from grimoirelab-metrics.md rename to grimoirelab-guide.md index cef72d7..457da98 100644 --- a/grimoirelab-metrics.md +++ b/grimoirelab-guide.md @@ -1,6 +1,6 @@ -# grimoirelab-metrics +# GrimoireLab Guide for HealthyCode -Client to generate GrimoireLab metrics for Project Health using the +Client to generate GrimoireLab metrics for project health using the software analytics platform [GrimoireLab](https://github.com/chaoss/grimoirelab). ![grimoirelab_metrics_schema.jpg](docs/images/grimoirelab_metrics_schema.jpg) From 4380ff9760ee23ed53f8e4c823ba52ca5aba356a Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 11:23:27 +0300 Subject: [PATCH 07/10] Update README.md fix minor typos and links --- README.md | 9 ++++----- 1 file changed, 4 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index 76c751a..843e748 100644 --- a/README.md +++ b/README.md @@ -59,8 +59,7 @@ Here is a trimmed result for the `semver` package: - Researchers studying open source sustainability ## When not to use it -HealthyCode measures project health. It does not scan for vulnerabilities or check license compliance. For those, see [VulnerableCode](https://github.com/aboutcode-org/vulnerablecode), [ScanCode.io](https://github.com/aboutcode-org/scancode.io), and -[ScanCode Toolkit](https://github.com/aboutcode-org/scancode-toolkit). +HealthyCode measures project health. It does not scan for vulnerabilities or check license compliance. For those, see [VulnerableCode](https://github.com/aboutcode-org/vulnerablecode), [ScanCode.io](https://github.com/aboutcode-org/scancode.io), and [ScanCode Toolkit](https://github.com/aboutcode-org/scancode-toolkit). ## Getting started HealthyCode needs a running [GrimoireLab 2.x](https://github.com/chaoss/grimoirelab/blob/2.x/README.md) @@ -83,7 +82,7 @@ docker run --rm ghcr.io/aboutcode-org/healthycode:0.2.0 \ Replace the URLs and credentials with your own. To analyze an SPDX SBOM instead of a single repository, mount the file into the container and pass its path. The first run for a repository takes longer because GrimoireLab has to collect its full history. -Installation, setup, and command-line usage guidance for GrimoireLab is in [grimoirelab-metrics.md]/(grimoirelab-metrics.md). +Installation, setup, and command-line usage guidance for GrimoireLab is in [grimoirelab-guide.md]/(grimoirelab-guide.md). ## How it works @@ -115,7 +114,7 @@ HealthyCode uses the Goal-Question-Metric (GQM) approach. You start with what yo GitHub stars are not on the list. They answer a question about popularity, and that is not what the model asks. ## The model -The v0.2 model was trained on about 1,150 popular npm packages drawn from Census II, Census III, and deps.dev. An open source expert reviewed 200 and classified 166 as healthy or unhealthy without seeing any scores. A logistic regression then learned which metrics best separate the two groups and how much weight each one gets. The data, notebook, a workflow diagram, and the expert's classifications guidelines are in [model/npm](model/npm). +The model was trained on about 1,150 popular npm packages drawn from Census II, Census III, and deps.dev. An open source expert reviewed 200 and classified 166 as healthy or unhealthy without seeing any scores. A logistic regression then learned which metrics best separate the two groups and how much weight each one gets. The data, notebook, a workflow diagram, and the expert's classifications guidelines are in [model/npm](model/npm). This is a first version. Weights and thresholds will be tuned as more packages and ecosystems are analyzed, and new questions and metrics will be added, drawing on [CHAOSS](https://www.chaoss.community/), [OpenSSF Scorecard](https://github.com/ossf/scorecard), and foundation maturity models. @@ -135,7 +134,7 @@ HealthyCode also fits into the AboutCode stack. You can run it as a ScanCode.io The full state-of-the-art review is available here: ## Project status -HealthyCode is at version 0.2.0 and under active development. npm is the first ecosystem, and the metrics and model will change as more packages are analyzed. If a result looks wrong to you, please open an issue with the package name and what you expected. That feedback goes straight into improving the model. +HealthyCode is under active development. npm is the first ecosystem, and the metrics and model will change as more packages are analyzed. If a result looks wrong to you, please open an issue with the package name and what you expected. That feedback goes straight into improving the model. ## Part of AboutCode HealthyCode is part of [AboutCode](https://aboutcode.org), a family of open source tools, open data, and open standards for healthy and secure software supply chains, alongside ScanCode, VulnerableCode, PurlDB, and Package-URL. It is developed by AboutCode and community contributors like you! From 3d0acf0a6bba89705c4a3284400d04e81bac48b8 Mon Sep 17 00:00:00 2001 From: adam Date: Sat, 3 Oct 2026 11:29:10 +0300 Subject: [PATCH 08/10] fix formatting minor changes to improve readability --- README.md | 22 +++++++++++----------- 1 file changed, 11 insertions(+), 11 deletions(-) diff --git a/README.md b/README.md index 843e748..d0686ad 100644 --- a/README.md +++ b/README.md @@ -111,8 +111,6 @@ HealthyCode uses the Goal-Question-Metric (GQM) approach. You start with what yo | --- | --- | --- | | The project stays maintained | Would it survive losing its top contributor? | Bus factor (CHAOSS Contributor Absence Factor)
Top author's share of commits over the last 12 months
Time since a second publisher shipped a release | -GitHub stars are not on the list. They answer a question about popularity, and that is not what the model asks. - ## The model The model was trained on about 1,150 popular npm packages drawn from Census II, Census III, and deps.dev. An open source expert reviewed 200 and classified 166 as healthy or unhealthy without seeing any scores. A logistic regression then learned which metrics best separate the two groups and how much weight each one gets. The data, notebook, a workflow diagram, and the expert's classifications guidelines are in [model/npm](model/npm). @@ -125,22 +123,24 @@ People have been working on open source health from three directions: 3. Community projects that work across ecosystems. CHAOSS defines open health metrics and models, and OpenSSF Scorecard scores security practices. Several commercial companies also sell package health scores. In most cases, you get the score but not the underlying data or the exact weights, which makes a result hard to check or reproduce. - -## HealthyCode is open code, open data, , every result can be audited -The code, the collected data, and the model weights are public. Each score comes with the metrics behind it and the settings used to produce it, and each metric comes from public commit history you can check yourself. You can also rerun the analysis on your own infrastructure to verify a result. Metrics follow CHAOSS definitions where they exist, and data collection is done by GrimoireLab, a CHAOSS project. -HealthyCode also fits into the AboutCode stack. You can run it as a ScanCode.io pipeline and look up results in PurlDB by PURL, the standard package identifier used in SBOMs and vulnerability databases and across software supply chains. That puts a package's health data next to its license and origin data from ScanCode for more comprehensive visibility into the packages you use. - The full state-of-the-art review is available here: - + +## With HealthyCode, every result can be audited +The code, the collected data, and the model weights are public. Each score comes with the metrics behind it and the settings used to produce it, and each metric comes from public commit history you can check yourself. You can also rerun the analysis on your own infrastructure to verify a result. Metrics follow CHAOSS definitions where they exist, and data collection is done by GrimoireLab, a CHAOSS project. + ## Project status -HealthyCode is under active development. npm is the first ecosystem, and the metrics and model will change as more packages are analyzed. If a result looks wrong to you, please open an issue with the package name and what you expected. That feedback goes straight into improving the model. +HealthyCode is under active development. npm is the first supported ecosystem, and the metrics and model will change as more packages are analyzed. If a result looks wrong to you, please open an issue with the package name and what you expected. That feedback goes straight into improving the model. ## Part of AboutCode -HealthyCode is part of [AboutCode](https://aboutcode.org), a family of open source tools, open data, and open standards for healthy and secure software supply chains, alongside ScanCode, VulnerableCode, PurlDB, and Package-URL. It is developed by AboutCode and community contributors like you! +HealthyCode is part of [AboutCode](https://aboutcode.org), a family of open source tools, open data, and open standards for healthy and secure software supply chains, alongside ScanCode, VulnerableCode, PurlDB, and Package-URL. + +You can run HealthyCode as a ScanCode.io pipeline and look up results in PurlDB by PURL, the standard package identifier used in SBOMs and vulnerability databases and across software supply chains. That puts a package's health data next to its license and origin data from ScanCode for more comprehensive visibility into the packages you use. + +HealthyCode is developed by AboutCode and community contributors like you! ## Contributing -Issues and pull requests are welcome. Please follow the [code of conduct](CODE_OF_CONDUCT.rst). +Issues and pull requests are welcome. Please follow the [Code of Conduct](CODE_OF_CONDUCT.rst). ## License The code is licensed under GPL-3.0-or-later. See [LICENSE](LICENSE). From e89156331d35524d325736d0afbc8f80c28547c8 Mon Sep 17 00:00:00 2001 From: adam Date: Sun, 4 Oct 2026 10:53:18 +0300 Subject: [PATCH 09/10] edited README.md fixed to clarify future plans, add images, and update links --- README.md | 41 +++++++++++++++++++++++++---------------- 1 file changed, 25 insertions(+), 16 deletions(-) diff --git a/README.md b/README.md index d0686ad..b42b432 100644 --- a/README.md +++ b/README.md @@ -62,9 +62,20 @@ Here is a trimmed result for the `semver` package: HealthyCode measures project health. It does not scan for vulnerabilities or check license compliance. For those, see [VulnerableCode](https://github.com/aboutcode-org/vulnerablecode), [ScanCode.io](https://github.com/aboutcode-org/scancode.io), and [ScanCode Toolkit](https://github.com/aboutcode-org/scancode-toolkit). ## Getting started -HealthyCode needs a running [GrimoireLab 2.x](https://github.com/chaoss/grimoirelab/blob/2.x/README.md) -instance with OpenSearch. GrimoireLab collects the data, and HealthyCode -turns it into metrics and a score. +Access health data by Package-URL (PURL). Health data is added to [PurlDB](https://github.com/aboutcode-org/purldb), accessible via API: https://health.purldb.io/api/ + +You ask for a package's health by its PURL (for example, `pkg:npm/semver`) and get the same JSON back. If the package has already been analyzed, the answer comes back right away. If not, the request starts a ScanCode.io pipeline to analyze it. Results will refresh when new versions are released. + +## Access health data through ScanCode.io pipelines +>This is planned future work. +Run HealthyCode from ScanCode.io, next to your other scans: https://github.com/aboutcode-org/scancode.io/blob/main/scanpipe/pipelines/scan_repo_health.py + +The `scan_repo_health` pipeline takes a Git repository URL, collects the data with GrimoireLab, and saves the metrics and score with the project results. Your ScanCode.io instance needs access to a GrimoireLab instance to run it. + +## Run HealthCode locally +>This is planned future work. + +HealthyCode needs a running [GrimoireLab 2.x](https://github.com/chaoss/grimoirelab/blob/2.x/README.md) instance with OpenSearch. GrimoireLab collects the data, and HealthyCode turns it into metrics and a score. Installation, setup, and command-line usage guidance for GrimoireLab is in [grimoirelab-guide.md]/(grimoirelab-guide.md). The quickest way to run HealthyCode is with the published Docker image: @@ -82,27 +93,23 @@ docker run --rm ghcr.io/aboutcode-org/healthycode:0.2.0 \ Replace the URLs and credentials with your own. To analyze an SPDX SBOM instead of a single repository, mount the file into the container and pass its path. The first run for a repository takes longer because GrimoireLab has to collect its full history. -Installation, setup, and command-line usage guidance for GrimoireLab is in [grimoirelab-guide.md]/(grimoirelab-guide.md). ## How it works - -1. You give HealthyCode a Git repository URL, or an SPDX SBOM that lists Git repositories. +1. HealthyCode is given a Git repository URL, or an SPDX SBOM that lists Git repositories. 2. HealthyCode asks GrimoireLab to collect each repository's history. Repositories GrimoireLab hasn't seen yet are added and analyzed. 3. When the data is ready, HealthyCode computes the metrics over a time window. The default is 12 months. 4. The npm health model turns those metrics into a score. 5. The metrics, score, and run settings are written to a JSON file. -HealthyCode runs on your own infrastructure. The only outside services it contacts are the public code hosts and registries the data comes from. +#### High-level overview of HealthyCode components +![High-level components of HealthyCode](https://github.com/user-attachments/assets/9fa54e45-a6f9-43b5-8c55-5d7159aa71ed) -## Access health data through ScanCode.io pipelines -You can run HealthyCode from ScanCode.io, next to your other scans: https://github.com/aboutcode-org/scancode.io/blob/main/scanpipe/pipelines/scan_repo_health.py - -The `scan_repo_health` pipeline takes a Git repository URL, collects the data with GrimoireLab, and saves the metrics and score with the project results. Your ScanCode.io instance needs access to a GrimoireLab instance to run it. +#### PurlDB and ScanCode.io overview +![PurlDB and ScanCode.io overview](https://github.com/user-attachments/assets/46d2c199-ddbc-43e0-8963-1a438bdcc20a) + +#### ScanCode.io + GrimoireLab overview +![ScanCode.io + GrimoireLab overview](https://github.com/user-attachments/assets/9d5d3bff-1e2d-4be3-8b86-132f02d9151e) -## Access health data by Package-URL (PURL) -We are adding health data to [PurlDB](https://github.com/aboutcode-org/purldb) via API: https://health.purldb.io/api/ - -You will be able to ask for a package's health by its PURL (for example, `pkg:npm/semver`) and get the same JSON back. If the package has already been analyzed, the answer comes back right away. If not, the request starts a ScanCode.io pipeline to analyze it. Results will refresh when new versions are released. ## The methodology HealthyCode uses the Goal-Question-Metric (GQM) approach. You start with what you care about (the goal), then work out what you need to know (the questions), and only then choose what you can measure (the metrics). For example: @@ -124,7 +131,7 @@ People have been working on open source health from three directions: Several commercial companies also sell package health scores. In most cases, you get the score but not the underlying data or the exact weights, which makes a result hard to check or reproduce. -The full state-of-the-art review is available here: +The full state-of-the-art review is available here: https://aboutcode.org/blog/npm-health-state-of-the-art ## With HealthyCode, every result can be audited The code, the collected data, and the model weights are public. Each score comes with the metrics behind it and the settings used to produce it, and each metric comes from public commit history you can check yourself. You can also rerun the analysis on your own infrastructure to verify a result. Metrics follow CHAOSS definitions where they exist, and data collection is done by GrimoireLab, a CHAOSS project. @@ -141,6 +148,8 @@ HealthyCode is developed by AboutCode and community contributors like you! ## Contributing Issues and pull requests are welcome. Please follow the [Code of Conduct](CODE_OF_CONDUCT.rst). + +Join the [AboutCode community Slack](https://join.slack.com/t/aboutcode-org/shared_invite/zt-31uzazd7l-tBHcqKUKkX6jUEPRLswiNw) to chat with maintainers, contributors, and users. ## License The code is licensed under GPL-3.0-or-later. See [LICENSE](LICENSE). From b3248f6fd2a240f623b8b8b421bd5e5fc7595d6c Mon Sep 17 00:00:00 2001 From: Philippe Ombredanne Date: Thu, 8 Oct 2026 00:59:59 +0200 Subject: [PATCH 10/10] Improve REDAME format and content Signed-off-by: Philippe Ombredanne --- README.md | 343 +++++++++++++++++++++++++++++++++++++++--------------- 1 file changed, 250 insertions(+), 93 deletions(-) diff --git a/README.md b/README.md index b42b432..d9922d8 100644 --- a/README.md +++ b/README.md @@ -1,100 +1,188 @@ # HealthyCode -Open source project health metrics and scores, built on an open model, open methodology, and open data. -HealthyCode tells you how healthy the open source projects you depend on are, so you can make informed, data-driven decisions about which components to adopt, upgrade, or replace. It reads a project's development history (commits, contributors, and how activity changes over time) and turns it into a set of health metrics and one health score. Every score comes with the metrics behind it, so you can see why a package scored the way it did and apply your own thresholds. +Open source project health metrics and scores, built on an open model, open +methodology, with open data. -HealthyCode uses [GrimoireLab](https://github.com/chaoss/grimoirelab) to collect data, [ScanCode.io](https://github.com/aboutcode-org/scancode.io) to run pipelines, and [PurlDB](https://github.com/aboutcode-org/purldb) to store the data. +HealthyCode attempts to tell you how healthy an open source projects is, so you +can make more informed, data-driven decisions about which packages to adopt, +upgrade, or replace. HealthyCode collect a project's development history events +(like commits, contributors, and other activities) and turns these into health +"metrics". The metrics are then used to compute a health score using a scoring +procedure where each metrics is weighted, and weights are tuned from a training +data sets. The score is backed by these metrics, enabling to see why a package +scored and how, and to set own policy thresholds. + +HealthyCode uses [GrimoireLab](https://github.com/chaoss/grimoirelab) to collect +data, [ScanCode.io](https://github.com/aboutcode-org/scancode.io) to run +pipelines, and [PurlDB](https://github.com/aboutcode-org/purldb) to store the +data. ## Why project health matters -Most software is assembled from open source packages. Security scanners are good at flagging a package with a known vulnerability, but they say nothing about a package whose last maintainer has quietly moved on... + +Most software is assembled from open source packages. Security scanners are good +at flagging a package with a known vulnerability, but they say nothing about a +package whose last maintainer has quietly moved on... > HealthyCode starts with npm. Other ecosystems will follow. > -> npm is the largest package registry in the world, and its packages depend on each other heavily. A small library can sit underneath thousands of projects. + +> npm is the largest package registry in the world, and its packages depend on +> each other heavily. A small library can sit underneath thousands of projects. > -> In npm, this risk spreads quickly. One unmaintained library can become a single point of failure for everything built on it, and it is also easier to take over. When a maintainer's account or email domain lapses, an attacker can claim it and publish a malicious release. + +> With npms, this risk can spread quickly as the JavaScript developers prefer +> publishing many smaller packages, and many small unmaintained library can become +> a single point of failure for everything built on it, and are also easier to +> take over. When a maintainer's account or email domain expires, an attacker can +> claim it and publish a malicious release. -HealthyCode helps you spot these packages before you depend on them, and keep watching the ones you already use. +HealthyCode helps you spot these packages before you depend on them, and helps +you keep watching the ones you already use as dependencies. + +## What's provided + +For this first iteration, HealthyCode focus is on npm packages. We will extend +support to other ecosystems, and we designed the scording model to be specific +to one open source packaging ecosystem. -## What you get -For each package, HealthyCode returns: -- Health metrics from the project's Git history, such as recent commits, contributor activity and growth, days since last commit, pony factor (how many people do half the work), and the elephant factor (how many organizations do half the work). The full list is in [METRICS.md](METRICS.md). -- A health score between 0 and 1 that is the model's estimate of how likely the project is to be unhealthy. 0 means healthy and 1 means unhealthy. -- The settings used for the run: time window, thresholds, and the versions of HealthyCode and the scoring model. +Given a PURL (Package-URL) for a package, HealthyCode's API returns: -Results are JSON, so you can feed them into a dashboard, a report, or an allow/block policy in your own tooling. HealthyCode gives you the evidence, but the policy decision stays with you. +- Health metrics from the project's Git history, such as recent commits, +contributor activity and growth, days since last commit, pony factor (how many +people do half the work), and the elephant factor (how many organizations do +half the work). The full list is in [METRICS.md](METRICS.md). -Here is a trimmed result for the `semver` package: +- A health score is calculted from the metrics, between 0 and 1 that is the +model's estimate of how likely the project is to be healthy. 0 means unhealthy +or risky, and 1 means healthier and less risky. + +- The settings used for the run: time window, thresholds, and the versions of +HealthyCode and the scoring model. + +The API results are JSON, so you can pull these into a dashboard, a report, or +in your software supply processing and management pipelines, including policies +to allow or block or alert in your own tooling about the health of a software +page. HealthyCode gives you the evidence, but the policy decision stays with +you. + +Here is a sample result for the `pkg:npm/semver` package: ```json { - "packages": { - "SPDXRef-Package-semver-954": { - "repository": "https://github.com/npm/node-semver.git", - "metrics": { - "recent_commits": 52, - "recent_contributors": 17, - "returning_contributors": 3, - "pony_factor": 2, - "elephant_factor": 1, - "days_since_last_commit": 1, - "found_file_license": 1 - }, - "score": { - "value": 0.0, - "metadata": { - "ecosystem": "npm", - "model": "health", - "version": "0.2" - } - } - } - } + "purl": "pkg:npm/semver", + "source_purl": "pkg:github/npm/node-semver", + "vcs_url": "https://github.com/npm/node-semver.git", + "scoring_model": "npm-health-0.2", + "score": 1.0, + "commit_range": { + "last_commit": "3484e1785e18a7ae4f06365f35c7b6eaec05088a", + "first_commit": "a25789b09b1192fa8414c35f2cd679ae2e1d5192", + "last_commit_date": "2026-09-10T19:51:52+00:00", + "first_commit_date": "2025-10-07T10:26:26-07:00" + }, + "run_start_date": "2026-10-07T21:22:06.088852Z", + "run_end_date": "2026-10-07T21:22:06.799342Z", + "metrics": { + "pony_factor": 2, + "total_commits": 45, + "recent_commits": 45, + "active_branches": 10, + "elephant_factor": 1, + "file_types_code": 37, + "commits_per_week": 0.8630136986301369, + "commits_per_year": 45.0, + "file_types_other": 61, + "commits_per_month": 3.6986301369863015, + "file_types_binary": 0, + "message_size_mean": 2538.8444444444444, + "contributor_growth": 5, + "found_file_license": 1, + "message_size_total": 114248, + "total_contributors": 16, + "found_file_adopters": 0, + "message_size_median": 593, + "recent_contributors": 16, + "total_organizations": 5, + "recent_organizations": 5, + "days_since_last_commit": 26, + "returning_contributors": 0, + "commit_size_added_lines": 930, + "contributor_growth_rate": 0.7142857142857143, + "coefficient_of_variation": 1.0786365421788087, + "commit_size_removed_lines": 137, + "commits_over_periods_rate": 1.0, + "developer_categories_core": 7, + "developer_categories_casual": 0, + "developer_categories_regular": 9, + "casual_regular_contributors_rate": 0.0 + }, + "date_collected": "2026-10-07T21:22:07.178202Z" } + ``` -## Who it's for -- Engineering teams deciding whether to adopt, upgrade, replace, or drop a dependency -- Open source program offices (OSPO) and supply chain teams tracking the packages their organization relies on +## Who is HealthyCode for? + +- Software and engineering teams deciding whether to adopt, upgrade, replace, or +drop a dependency in their codebases + +- Open source program offices (OSPO) and software supply chains teams tracking +the packages used in their organization apps, systems and products + - Maintainers who want an outside view of their own project + - Researchers studying open source sustainability -## When not to use it -HealthyCode measures project health. It does not scan for vulnerabilities or check license compliance. For those, see [VulnerableCode](https://github.com/aboutcode-org/vulnerablecode), [ScanCode.io](https://github.com/aboutcode-org/scancode.io), and [ScanCode Toolkit](https://github.com/aboutcode-org/scancode-toolkit). +## When not to use HealthyCode + +HealthyCode measures or rather approximates an open source project's health. It +does not scan for vulnerabilities or check license compliance or check the +configuration or security posture of a project. For those, see: + +- [OpenSSF ScoreCard](https://github.com/ossf/scorecard) for security posture and configuration +- [VulnerableCode](https://github.com/aboutcode-org/vulnerablecode) for vulnerability lookup +- [ScanCode.io](https://github.com/aboutcode-org/scancode.io) to orchestrate scans including health scans +- [ScanCode Toolkit](https://github.com/aboutcode-org/scancode-toolkit) for origin, license, copyright and dependencies +- [PurlDB](https://github.com/aboutcode-org/purldb) that also hosts the health/ API endpoint. +- [ClearlyDefined](https://https://github.com/clearlydefined/) that pre-scans, and curates open source packages. + ## Getting started -Access health data by Package-URL (PURL). Health data is added to [PurlDB](https://github.com/aboutcode-org/purldb), accessible via API: https://health.purldb.io/api/ + +You access health data by Package-URL (PURL) using the [PurlDB](https://github.com/aboutcode-org/purldb) /health API. +This is accessible at https://health.purldb.io/api/health for demonstration. +For instance, check https://health.purldb.io/api/health/?purl=pkg:npm/semver or https://health.purldb.io/api/health/?purl=pkg:npm/lodash -You ask for a package's health by its PURL (for example, `pkg:npm/semver`) and get the same JSON back. If the package has already been analyzed, the answer comes back right away. If not, the request starts a ScanCode.io pipeline to analyze it. Results will refresh when new versions are released. +You call the health API for a package using its PURL (for example, +`pkg:npm/semver`) and get JSON back. If the package has already been analyzed, +the answer comes back immediately. If not, you can poll the same URL until the +metrics and scores are collected and computed. The request will return first a +"new" status and then a "submitted" status when the job is processed. +The results are cached for a week, and the JSON is returned last. + +The processing goes from PurlDB `/health` api endpoint through ScanCode.io +`scan_repo_health` pipeline to the HealthyCode npm-health scoring model (for +now, for npms only), that calls GrimoireLab to collect project metrics. ## Access health data through ScanCode.io pipelines ->This is planned future work. -Run HealthyCode from ScanCode.io, next to your other scans: https://github.com/aboutcode-org/scancode.io/blob/main/scanpipe/pipelines/scan_repo_health.py - -The `scan_repo_health` pipeline takes a Git repository URL, collects the data with GrimoireLab, and saves the metrics and score with the project results. Your ScanCode.io instance needs access to a GrimoireLab instance to run it. + +Run HealthyCode using ScanCode.io `scan_repo_health` next to your other scans: +https://github.com/aboutcode-org/scancode.io/blob/main/scanpipe/pipelines/scan_repo_health.py + +That pipeline takes a Git repository URL, collects the data with GrimoireLab, +and saves the metrics and score with the project results. Your ScanCode.io needs +access to a configured GrimoireLab instance and HealthCode models. ## Run HealthCode locally ->This is planned future work. - -HealthyCode needs a running [GrimoireLab 2.x](https://github.com/chaoss/grimoirelab/blob/2.x/README.md) instance with OpenSearch. GrimoireLab collects the data, and HealthyCode turns it into metrics and a score. Installation, setup, and command-line usage guidance for GrimoireLab is in [grimoirelab-guide.md]/(grimoirelab-guide.md). - -The quickest way to run HealthyCode is with the published Docker image: - -```bash -docker run --rm ghcr.io/aboutcode-org/healthycode:0.2.0 \ - /opt/healthycode/.venv/bin/grimoirelab-metrics \ - https://github.com/npm/node-semver.git \ - --grimoirelab-url http://your-grimoirelab:8000 \ - --grimoirelab-user USER --grimoirelab-password PASSWORD \ - --grimoirelab-ecosystem npm --grimoirelab-project my-packages \ - --opensearch-url https://your-opensearch:9200 \ - --opensearch-user USER --opensearch-password PASSWORD \ - > metrics.json -``` -Replace the URLs and credentials with your own. To analyze an SPDX SBOM instead of a single repository, mount the file into the container and pass its path. The first run for a repository takes longer because GrimoireLab has to collect its full history. +HealthyCode needs a running [GrimoireLab 2.x](https://github.com/chaoss/grimoirelab/blob/2.x/README.md) with +OpenSearch. GrimoireLab collects the data, and HealthyCode turns the data into metrics +(like the menagerie of ponies and elephants) and computes a score. +The installation, setup, and command-line usage guidance for GrimoireLab is in +[grimoirelab-guide.md]/(grimoirelab-guide.md). -## How it works + +## How Healthy works 1. HealthyCode is given a Git repository URL, or an SPDX SBOM that lists Git repositories. 2. HealthyCode asks GrimoireLab to collect each repository's history. Repositories GrimoireLab hasn't seen yet are added and analyzed. 3. When the data is ready, HealthyCode computes the metrics over a time window. The default is 12 months. @@ -102,56 +190,125 @@ Replace the URLs and credentials with your own. To analyze an SPDX SBOM instead 5. The metrics, score, and run settings are written to a JSON file. #### High-level overview of HealthyCode components + ![High-level components of HealthyCode](https://github.com/user-attachments/assets/9fa54e45-a6f9-43b5-8c55-5d7159aa71ed) -#### PurlDB and ScanCode.io overview +#### Drilling down on PurlDB and ScanCode.io + ![PurlDB and ScanCode.io overview](https://github.com/user-attachments/assets/46d2c199-ddbc-43e0-8963-1a438bdcc20a) -#### ScanCode.io + GrimoireLab overview +#### Drilling down on ScanCode.io + GrimoireLab + ![ScanCode.io + GrimoireLab overview](https://github.com/user-attachments/assets/9d5d3bff-1e2d-4be3-8b86-132f02d9151e) -## The methodology -HealthyCode uses the Goal-Question-Metric (GQM) approach. You start with what you care about (the goal), then work out what you need to know (the questions), and only then choose what you can measure (the metrics). For example: +## Methodology for selecting metrics -| Goal | Question | Metrics | -| --- | --- | --- | -| The project stays maintained | Would it survive losing its top contributor? | Bus factor (CHAOSS Contributor Absence Factor)
Top author's share of commits over the last 12 months
Time since a second publisher shipped a release | +HealthyCode uses the Goal-Question-Metric (GQM) approach. You start with what +you care about (the goal), then work out what you need to know (the questions), +and only then choose what you can measure (the metrics). -## The model -The model was trained on about 1,150 popular npm packages drawn from Census II, Census III, and deps.dev. An open source expert reviewed 200 and classified 166 as healthy or unhealthy without seeing any scores. A logistic regression then learned which metrics best separate the two groups and how much weight each one gets. The data, notebook, a workflow diagram, and the expert's classifications guidelines are in [model/npm](model/npm). +For example: -This is a first version. Weights and thresholds will be tuned as more packages and ecosystems are analyzed, and new questions and metrics will be added, drawing on [CHAOSS](https://www.chaoss.community/), [OpenSSF Scorecard](https://github.com/ossf/scorecard), and foundation maturity models. +- Goal: The project stays maintained +- Question: Would it survive losing its top contributor? +- Metrics: Pony factor (CHAOSS Contributor Absence Factor aka. "Bus Factor" ), +e.g., top author's share of commits over the last 12 months. -## How HealthyCode compares -People have been working on open source health from three directions: -1. Organizations improving how they take in open source, through work such as OpenChain, the FINOS Open Source Maturity Model, the Good Governance Initiative, and the TODO Group. -2. Foundations grading their own projects, such as the Apache Software Foundation maturity model, the Eclipse Foundation Development Process, and the CNCF project lifecycle (sandbox, incubating, graduated, archived). -3. Community projects that work across ecosystems. CHAOSS defines open health metrics and models, and OpenSSF Scorecard scores security practices. +## (npm-health) scoring model -Several commercial companies also sell package health scores. In most cases, you get the score but not the underlying data or the exact weights, which makes a result hard to check or reproduce. - -The full state-of-the-art review is available here: https://aboutcode.org/blog/npm-health-state-of-the-art +The model was trained on about 1,150 popular npm packages drawn from Census II, +Census III, and deps.dev. As open source experts, we reviewed about 200 of these +npms and classified them as "healthy" or "unhealthy" to create a training +dataset. We computed a logistic regression to determine and learn learned which +metrics best dsicriminate between the two groups and how much weight each metric +would get in the score computation. + +All the backing data, the Jupyter notebook, a workflow diagram, and our expert +classifications guidelines are in the [model/npm](model/npm) directory for +reference. + +This is a first version and the weights and thresholds will be tuned as more +packages are analyzed; we will define new models fo other ecosystems using the +same approch; and new questions and metrics will be added, drawing on +[CHAOSS](https://www.chaoss.community/), +[OpenSSF Scorecard](https://github.com/ossf/scorecard), +and foundation maturity models. + +Eventually we plan to enrich back the [OpenSSF +Scorecard](https://github.com/ossf/scorecard) with these metrics as additionals +checks. + + +## How this project compares to other approaches and alterives + +We have published a review of the state-of-the-art here: +https://aboutcode.org/blog/npm-health-state-of-the-art + +People have been working on open source health from these three directions: + +1. Organizations improving how they reuse and contribute to open source, through +work such as OpenChain, the FINOS Open Source Maturity Model, the Good +Governance Initiative, and the TODO Group. + +2. Open source Foundations grading their own projects, such as the Apache +Software Foundation maturity model, the Eclipse Foundation Development Process, +and the CNCF project lifecycle. + +3. Community projects that work across ecosystems, such as the CHAOSS project +that defines open health metrics and models (and where the GrimoireLab projects +lives), and the OpenSSF Scorecard that scores security practices, project +configuration and security posture. + +Multiple commercial companies also sell or publish open source package health +scores, but in most cases these are opaque, and you get the score without the +underlying data or weight calculation, making results hard to review, trace or +reproduce. -## With HealthyCode, every result can be audited -The code, the collected data, and the model weights are public. Each score comes with the metrics behind it and the settings used to produce it, and each metric comes from public commit history you can check yourself. You can also rerun the analysis on your own infrastructure to verify a result. Metrics follow CHAOSS definitions where they exist, and data collection is done by GrimoireLab, a CHAOSS project. +## Most importantly, our results can be audited and traced + +The code, the collected data, and the model weights are open and public. Each +score is backed and comed by its supporting metrics; metrics comes from public +project data and history that can be verified. + +And you can also rerun the analysis on your own infrastructure to verify, and +reproduce the results. The metrics follow CHAOSS definitions where they exist, +and data collection is run by GrimoireLab, a CHAOSS project. ## Project status -HealthyCode is under active development. npm is the first supported ecosystem, and the metrics and model will change as more packages are analyzed. If a result looks wrong to you, please open an issue with the package name and what you expected. That feedback goes straight into improving the model. -## Part of AboutCode -HealthyCode is part of [AboutCode](https://aboutcode.org), a family of open source tools, open data, and open standards for healthy and secure software supply chains, alongside ScanCode, VulnerableCode, PurlDB, and Package-URL. +HealthyCode is under active development with npm as the first supported +ecosystem. The selected metrics and models will change as more packages are +analyzed, and if a result looks wrong to you, please open an issue with the +Package-URL and a hint of what you expected. That feedback will go directlyinto +improving the model with new and improved training data! + +## HealthyCode in the larger AboutCode context + +HealthyCode is an initiative that is part of [AboutCode](https://aboutcode.org), +a family of open source tools, open data, and open standards for healthy and +safe software supply chains, together with ScanCode, VulnerableCode, PurlDB, +DejaCode, and Package-URL. -You can run HealthyCode as a ScanCode.io pipeline and look up results in PurlDB by PURL, the standard package identifier used in SBOMs and vulnerability databases and across software supply chains. That puts a package's health data next to its license and origin data from ScanCode for more comprehensive visibility into the packages you use. +You can get results in the PurlDB health API by PURL, the standard package +identifier used in SBOMs and vulnerability databases and across software supply +chains. Exposing the /health endpoint in PurlDB also puts a package's health +data next to its license and origin data collected in PurlDB from ScanCode, +eventually delivering better visibility into the packages from multiple angles. -HealthyCode is developed by AboutCode and community contributors like you! +HealthyCode is developed by the AboutCode and GrimoireLab community and +community contributors like you! ## Contributing -Issues and pull requests are welcome. Please follow the [Code of Conduct](CODE_OF_CONDUCT.rst). + +Issues and pull requests are welcome. +Please follow the [Code of Conduct](CODE_OF_CONDUCT.rst). -Join the [AboutCode community Slack](https://join.slack.com/t/aboutcode-org/shared_invite/zt-31uzazd7l-tBHcqKUKkX6jUEPRLswiNw) to chat with maintainers, contributors, and users. +Join the [AboutCode community Slack](https://join.slack.com/t/aboutcode-org/shared_invite/zt-31uzazd7l-tBHcqKUKkX6jUEPRLswiNw) +to chat with maintainers, contributors, and users. ## License + The code is licensed under GPL-3.0-or-later. See [LICENSE](LICENSE). Data produced by HealthyCode is licensed under