Skip to content
GenentechPublic

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

22 Commits

Folders and files

Repository files navigation

CHART

CHART (Cellular High-content imaging Archetype Response Toolkit) translates high-content imaging features into interpretable morphological phenotypes. CHART identifies marker-specific cellular archetypes, annotates representative images using multimodal large language models, and integrates co-regulated archetypes into morphological programs that describe coordinated cellular responses to perturbation.

Overview of the CHART framework

Components

  • OPS preprocessing — quality control, filtering, normalization, outlier detection, guide filtering and single-cell/aggregated embeddings.
  • Compartment archetype analysis — identification of distinctive reference morphological states for each imaged marker.
  • Automated archetype annotation — contrastive image analysis and standardized reports generated with multimodal language models and user review.
  • Morphological program definition — integration of co-regulated archetypes into interpretable programs of coordinated cellular change.

Requirements

Most of CHART is Python and arrives with the package. The archetype analysis is the exception: every step of it but the first is an R script that chart archetypes runs through Rscript, so that stage needs R and eight R packages. environment.yml describes them:

mamba env create -f environment.yml
conda activate chart-r
Rscript -e 'install.packages("archetypes", repos="https://cloud.r-project.org")'

Quickstart

Install the package, plus the environment above if you want the archetype analysis:

pip install .

Write a config saying where the input is and where output should go:

data_output_dir: /work/screen1/bulk
local_output_dir: /work/screen1/reports

preprocessing_bywell:
  merged_dir: /data/screen1/objects        # where the input parquet lives

archetypes:
  channels: [TOM20]                        # the channels to analyse
  all_channels: [DAPI1, TOM20, Golgin97]   # every channel the table holds

Then run the two stages in order:

chart preprocessing --config my_screen.yaml --all-wells
chart archetypes --config my_screen.yaml --all-channels

Wells are read off the input filenames. Name wells individually with --well and channels with --channel, and use --steps on either command to run only part of a pipeline. chart archetypes --dry-run prints every command the stage would run without running any of them.

Configuration

Every step reads one YAML file, given with --config:

chart preprocessing --config my_screen.yaml --all-wells

Use example_config.yaml as a starting point: copy it and edit the paths. Every key in it except the two output directories is optional, and the values shown are the defaults.

The schema block maps CHART variables to the input screen, accounting for differences in input column names and feature naming conventions:

schema:
  cell_column: object_number
  guide_column: sgrna_id
  gene_column: target_gene
  feature_patterns: ['^Nuclei_', '^Cells_', '^Cytoplasm_']
  control_gene: safe_harbour
  control_prefix: null
  centroid_columns: [Nuclei_Location_Center_Y, Nuclei_Location_Center_X]

Every column in feature_patterns must be numeric.

Input filenames are configurable in the same file, under preprocessing_bywell. Only .parquet input files are accepted.

preprocessing_bywell:
  objects_pattern: 'plate1_{well}.objects.parquet'
  features_pattern: 'plate1_{well}.features.parquet'

When one file per well holds objects and features together, set
premerged: true and name it with merged_pattern instead.

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages