scgpt

GitHub

利用scGPT单细胞基础模型对单细胞表达数据进行嵌入、聚类、整合及细胞类型注释。支持零样本或微调,并提供基因级表示以用于扰动和GRN任务,适用于生物信息学数据分析流程。

resources/skills/scgpt/SKILL.md aipoch/open-science

Trigger Scenarios

需要生成单细胞数据的嵌入向量 执行单细胞类型的自动注释 进行基因水平的特征表示分析

Install

npx skills add aipoch/open-science --skill scgpt -g -y
More Options

Non-standard path

npx skills add https://github.com/aipoch/open-science/tree/main/resources/skills/scgpt -g -y

Use without installing

npx skills use aipoch/open-science@scgpt

指定 Agent (Claude Code)

npx skills add aipoch/open-science --skill scgpt -a claude-code -g -y

安装 repo 全部 skill

npx skills add aipoch/open-science --all -g -y

预览 repo 内 skill

npx skills add aipoch/open-science --list

SKILL.md

Frontmatter
{
    "name": "scgpt",
    "license": "Apache-2.0",
    "category": "biomodels",
    "metadata": {
        "third_party": [
            {
                "kind": "weights",
                "name": "scGPT",
                "info_url": "https:\/\/github.com\/bowang-lab\/scGPT",
                "provider": "Wang Lab (University of Toronto)"
            }
        ],
        "display-name": "scGPT"
    },
    "description": "Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering\/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation\/GRN tasks.\nFor probabilistic single-cell models (scVI etc.), use the scvi-tools library.\n",
    "requirements": [
        "gpu"
    ]
}

scGPT — Single-Cell Foundation Model

Prerequisites

Requirement Minimum Recommended
Python 3.10+ 3.11
CUDA 12.1+ 12.4+
GPU VRAM 16 GB 24 GB+

How to run

Loading the vocabulary and checkpoint

scGPT checkpoints are raw directories (args.json, best_model.pt, vocab.json) — not Hugging Face hub repos. Point at the directory, not an HF repo id.

from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv))   # 60697 for the released human checkpoint

Embedding an AnnData

import anndata as ad
from scgpt.tasks import embed_data

adata = ad.read_h5ad("dataset.h5ad")        # var must contain a gene-name column
emb = embed_data(
    adata,
    model_dir="/path/to/scgpt-human",
    gene_col="feature_name",
    use_fast_transformer=False,             # see Gotchas
)
# emb is an AnnData with .obsm["X_scGPT"]

Output format

embed_data returns an AnnData whose .obsm["X_scGPT"] is the per-cell embedding (n_cells × emb_dim, 512 by default). Downstream: feed to scanpy.pp.neighbors / scanpy.tl.umap.

Remote compute

Needs ≥24 GB VRAM and the released human checkpoint (~200 MB: args.json, best_model.pt, vocab.json). Read compute_details({provider, mode:'read'}) for an environment with scgpt and a pre-cached checkpoint directory, then:

c = host.compute.create(provider)
job = c.submitJob(
    intent="scGPT embed 50k cells — 1×GPU, ~5 min",
    inputs=[
        {"src": "dataset.h5ad", "dstFilename": "dataset.h5ad"},
        {"src": "embed.py", "dstFilename": "embed.py"},
    ],
    command="python3 embed.py",
    environment=...,   # env name from compute_details
    outputs=["embedded.h5ad"],
    timeoutSeconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Retain the exact returned job_id. Query that saved ID with the non-blocking host.compute.create(provider).attachJob(job_id).status() or .result() when its state or result is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed or had already been committed; otherwise the app starts the later analysis turn for an unread final result. See the remote-compute-ssh skill for the orchestration details.

See the remote-compute-ssh / remote-compute-modal skill for the orchestration details.

In embed.py, pass model_dir= the checkpoint path from compute_details. If flash-attn is unavailable in that environment, set use_fast_transformer=False.

Gotchas

  • use_fast_transformer default is True but resolves to a FlashAttention path that may not import in every env. Pass use_fast_transformer=False unless you've confirmed flash_attn loads cleanly.
  • The package historically depended on torchtext.vocab.Vocab; in environments without torchtext a pure-Python shim provides Vocab — functionally identical for GeneVocab, but if you hit AttributeError: 'Vocab' object has no attribute …, you're on a stale shim.
  • Gene names must match the vocab; unmatched genes are dropped. Set gene_col to the column in adata.var that holds symbols.

Troubleshooting

Symptom Fix
flash_attn is not installed warning at import Harmless; pass use_fast_transformer=False
'Vocab' object has no attribute 'vocab' Env has an old torchtext shim — update the env
Nearly all genes dropped Wrong gene_col; check adata.var.columns
"scgpt not in manifest" / env-detection misses scGPT The baked env manifest lists the distribution as scGPT (and flash_attn), pip's canonical casing — normalize manifest keys before lookup: name.lower().replace('-', '_')

Next: cluster/annotate the embedding with the scanpy library (sc.pp.neighborssc.tl.leiden / sc.tl.umap), or compare to an scvi-tools latent space on the same data.

Version History

  • 44394f0 Current 2026-09-11 11:15

Same Skill Collection

resources/skills/alphafold2/SKILL.md
resources/skills/boltz/SKILL.md
resources/skills/borzoi/SKILL.md
resources/skills/chai1/SKILL.md
resources/skills/compute-env-setup/SKILL.md
resources/skills/customize/SKILL.md
resources/skills/diffdock/SKILL.md
resources/skills/env-management/SKILL.md
resources/skills/evo2/SKILL.md
resources/skills/fair-esm2/SKILL.md
resources/skills/figure-composer/SKILL.md
resources/skills/figure-style/SKILL.md
resources/skills/indication-dossier/SKILL.md
resources/skills/ligandmpnn/SKILL.md
resources/skills/literature-review/SKILL.md
resources/skills/openfold3/SKILL.md
resources/skills/paper-narrative/SKILL.md
resources/skills/proteinmpnn/SKILL.md
resources/skills/remote-compute-ssh/SKILL.md
resources/skills/scvi-tools/SKILL.md
resources/skills/self-awareness/SKILL.md
resources/skills/skill-creator/SKILL.md
resources/skills/solublempnn/SKILL.md
resources/skills/esmfold2/SKILL.md

Metadata

Files
0
Version
44394f0
Hash
a87b3b7f
Indexed
2026-09-11 11:15

ホーム - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-12 07:25
浙ICP备14020137号-1 $お客様$