scgpt
GitHub利用scGPT单细胞基础模型对单细胞表达数据进行嵌入、聚类、整合及细胞类型注释。支持零样本或微调,并提供基因级表示以用于扰动和GRN任务,适用于生物信息学数据分析流程。
Trigger Scenarios
Install
npx skills add aipoch/open-science --skill scgpt -g -y
SKILL.md
Frontmatter
{
"name": "scgpt",
"license": "Apache-2.0",
"category": "biomodels",
"metadata": {
"third_party": [
{
"kind": "weights",
"name": "scGPT",
"info_url": "https:\/\/github.com\/bowang-lab\/scGPT",
"provider": "Wang Lab (University of Toronto)"
}
],
"display-name": "scGPT"
},
"description": "Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering\/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation\/GRN tasks.\nFor probabilistic single-cell models (scVI etc.), use the scvi-tools library.\n",
"requirements": [
"gpu"
]
}
scGPT — Single-Cell Foundation Model
Prerequisites
| Requirement | Minimum | Recommended |
|---|---|---|
| Python | 3.10+ | 3.11 |
| CUDA | 12.1+ | 12.4+ |
| GPU VRAM | 16 GB | 24 GB+ |
How to run
Loading the vocabulary and checkpoint
scGPT checkpoints are raw directories (args.json, best_model.pt,
vocab.json) — not Hugging Face hub repos. Point at the directory, not an HF
repo id.
from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv)) # 60697 for the released human checkpoint
Embedding an AnnData
import anndata as ad
from scgpt.tasks import embed_data
adata = ad.read_h5ad("dataset.h5ad") # var must contain a gene-name column
emb = embed_data(
adata,
model_dir="/path/to/scgpt-human",
gene_col="feature_name",
use_fast_transformer=False, # see Gotchas
)
# emb is an AnnData with .obsm["X_scGPT"]
Output format
embed_data returns an AnnData whose .obsm["X_scGPT"] is the per-cell
embedding (n_cells × emb_dim, 512 by default). Downstream: feed to
scanpy.pp.neighbors / scanpy.tl.umap.
Remote compute
Needs ≥24 GB VRAM and the released human checkpoint (~200 MB:
args.json, best_model.pt, vocab.json). Read
compute_details({provider, mode:'read'}) for an environment with scgpt
and a pre-cached checkpoint directory, then:
c = host.compute.create(provider)
job = c.submitJob(
intent="scGPT embed 50k cells — 1×GPU, ~5 min",
inputs=[
{"src": "dataset.h5ad", "dstFilename": "dataset.h5ad"},
{"src": "embed.py", "dstFilename": "embed.py"},
],
command="python3 embed.py",
environment=..., # env name from compute_details
outputs=["embedded.h5ad"],
timeoutSeconds=1800,
)
print(job.job_id) # cell ends here — kernel never blocks on compute
Retain the exact returned job_id. Query that saved ID with the non-blocking
host.compute.create(provider).attachJob(job_id).status() or .result() when its state or result
is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed
or had already been committed; otherwise the app starts the later analysis turn for an unread
final result. See the remote-compute-ssh skill for the orchestration details.
See the remote-compute-ssh / remote-compute-modal skill for the
orchestration details.
In embed.py, pass model_dir= the checkpoint path from compute_details.
If flash-attn is unavailable in that environment, set
use_fast_transformer=False.
Gotchas
use_fast_transformerdefault isTruebut resolves to a FlashAttention path that may not import in every env. Passuse_fast_transformer=Falseunless you've confirmedflash_attnloads cleanly.- The package historically depended on
torchtext.vocab.Vocab; in environments without torchtext a pure-Python shim providesVocab— functionally identical forGeneVocab, but if you hitAttributeError: 'Vocab' object has no attribute …, you're on a stale shim. - Gene names must match the vocab; unmatched genes are dropped. Set
gene_colto the column inadata.varthat holds symbols.
Troubleshooting
| Symptom | Fix |
|---|---|
flash_attn is not installed warning at import |
Harmless; pass use_fast_transformer=False |
'Vocab' object has no attribute 'vocab' |
Env has an old torchtext shim — update the env |
| Nearly all genes dropped | Wrong gene_col; check adata.var.columns |
| "scgpt not in manifest" / env-detection misses scGPT | The baked env manifest lists the distribution as scGPT (and flash_attn), pip's canonical casing — normalize manifest keys before lookup: name.lower().replace('-', '_') |
Next: cluster/annotate the embedding with the scanpy library
(sc.pp.neighbors → sc.tl.leiden / sc.tl.umap), or compare to an
scvi-tools latent space on the same data.
Version History
- 44394f0 Current 2026-09-11 11:15


