CHART (Cellular High-content imaging Archetype Response Toolkit) translates high-content imaging features into interpretable morphological phenotypes. CHART identifies marker-specific cellular archetypes, annotates representative images using multimodal large language models, and integrates co-regulated archetypes into morphological programs that describe coordinated cellular responses to perturbation.
- OPS preprocessing — quality control, filtering, normalization, outlier detection, guide filtering and single-cell/aggregated embeddings.
- Compartment archetype analysis — identification of distinctive reference morphological states for each imaged marker.
- Automated archetype annotation — contrastive image analysis and standardized reports generated with multimodal language models and user review.
- Morphological program definition — integration of co-regulated archetypes into interpretable programs of coordinated cellular change.
Most of CHART is Python and arrives with the package. The archetype
analysis is the exception: every step of it but the first is an R script
that chart archetypes runs through Rscript, so that stage needs R and
eight R packages. environment.yml describes them:
mamba env create -f environment.yml
conda activate chart-r
Rscript -e 'install.packages("archetypes", repos="https://cloud.r-project.org")'
Install the package, plus the environment above if you want the archetype analysis:
pip install .
Write a config saying where the input is and where output should go:
data_output_dir: /work/screen1/bulk
local_output_dir: /work/screen1/reports
preprocessing_bywell:
merged_dir: /data/screen1/objects # where the input parquet lives
archetypes:
channels: [TOM20] # the channels to analyse
all_channels: [DAPI1, TOM20, Golgin97] # every channel the table holdsThen run the two stages in order:
chart preprocessing --config my_screen.yaml --all-wells
chart archetypes --config my_screen.yaml --all-channels
Wells are read off the input filenames. Name wells
individually with --well and channels with --channel, and use
--steps on either command to run only part of a pipeline.
chart archetypes --dry-run prints every command the stage would run
without running any of them.
Every step reads one YAML file, given with --config:
chart preprocessing --config my_screen.yaml --all-wells
Use example_config.yaml as a starting point: copy it and edit the paths. Every key in it except the two output directories is optional, and the values shown are the defaults.
The schema block maps CHART variables to the input screen, accounting for differences in input column names and feature naming conventions:
schema:
cell_column: object_number
guide_column: sgrna_id
gene_column: target_gene
feature_patterns: ['^Nuclei_', '^Cells_', '^Cytoplasm_']
control_gene: safe_harbour
control_prefix: null
centroid_columns: [Nuclei_Location_Center_Y, Nuclei_Location_Center_X]Every column in feature_patterns must be numeric.
Input filenames are configurable in the same file, under
preprocessing_bywell. Only .parquet input files are accepted.
preprocessing_bywell:
objects_pattern: 'plate1_{well}.objects.parquet'
features_pattern: 'plate1_{well}.features.parquet'When one file per well holds objects and features together, set
premerged: true and name it with merged_pattern instead.
