evo2

GitHub

基于Evo2基因组大模型,提供DNA序列打分、嵌入及生成能力,用于变异效应评估和调控区域分析。

resources/skills/evo2/SKILL.md aipoch/open-science

Trigger Scenarios

计算DNA序列的对数似然概率 生成或编辑DNA序列 进行基因组变异的效应评分

Install

npx skills add aipoch/open-science --skill evo2 -g -y
More Options

Non-standard path

npx skills add https://github.com/aipoch/open-science/tree/main/resources/skills/evo2 -g -y

Use without installing

npx skills use aipoch/open-science@evo2

指定 Agent (Claude Code)

npx skills add aipoch/open-science --skill evo2 -a claude-code -g -y

安装 repo 全部 skill

npx skills add aipoch/open-science --all -g -y

预览 repo 内 skill

npx skills add aipoch/open-science --list

SKILL.md

Frontmatter
{
    "name": "evo2",
    "license": "Apache-2.0",
    "category": "biomodels",
    "metadata": {
        "third_party": [
            {
                "kind": "weights",
                "name": "Evo 2",
                "license": "Apache-2.0",
                "provider": "Arc Institute",
                "terms_url": "https:\/\/github.com\/ArcInstitute\/evo2\/blob\/main\/LICENSE"
            }
        ],
        "display-name": "Evo 2"
    },
    "description": "Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model. Use this skill when: (1) Computing per-nucleotide or per-sequence likelihoods for variant effect\n    scoring,\n(2) Embedding genomic windows for downstream classification, (3) Generating DNA conditioned on a prefix, (4) Scoring regulatory or coding regions across species.\n",
    "requirements": [
        "gpu"
    ]
}

Evo 2 — DNA Language Model

Prerequisites

Requirement Minimum Recommended
Python 3.11 3.12 (<3.13)
CUDA 12.1+ 12.4+
GPU VRAM 24 GB (7B bf16) 80 GB (40B)
RAM 32 GB 128 GB

How to run

Installation

pip install evo2
# Weights pulled from Hugging Face on first model load.

Loading and scoring

from evo2 import Evo2

model = Evo2("evo2_7b")        # or "evo2_40b" — see model table
seqs = ["ATCG" * 50, "GGGCTTAA" * 25]
ll = model.score_sequences(seqs)   # → list[float], mean per-token log-likelihood
print(ll)

Generation

out = model.generate(
    prompt_seqs=["ATGAAAGCT"],
    n_tokens=256,
    temperature=0.7,
)
print(out.sequences[0])

Models

Name Params Context VRAM (bf16) Notes
evo2_7b 7 B 1 M nt ~22 GB Default; fits on a single 24 GB+ GPU
evo2_40b 40 B 1 M nt ~78 GB H100 80 GB or multi-GPU
evo2_1b_base 1 B 8 K nt ~6 GB FP8 path requires sm_89+ (H100)

Output format

score_sequences returns a list[float] (or np.ndarray) of mean log-likelihoods, one per input sequence. More negative ⇒ less likely under the model. For variant effect, compute Δll = ll_alt - ll_ref over a fixed window.

generate returns a GenerationOutput with .sequences (list[str]), .logits (list[Tensor]), and .logprobs_mean (list[float]) — always populated, no flag required.

Decision tree

Need a DNA model?
│
├─ Per-base/per-sequence likelihood, generation → Evo 2 ✓
├─ Predict experimental tracks (expression, accessibility) → borzoi
└─ Protein, not DNA → fair-esm2 / esmfold2

Remote compute

7B/40B inference is GPU-bound (≥24 GB / 80 GB VRAM). Read compute_details({provider, mode:'read'}) for an environment with evo2 + flash-attn and a pre-cached HF weight mount, then submit:

c = host.compute.create(provider)
job = c.submitJob(
    intent="Evo2-7B score 200bp variant window — 1×GPU, ~2 min",
    inputs=[{"src": "score_evo2.py", "dstFilename": "score_evo2.py"}],
    command="python3 score_evo2.py",   # env selection is host-specific — see compute_details for your provider
    outputs=["scores.json"],
    timeoutSeconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Retain the exact returned job_id. Query that saved ID with the non-blocking c.attachJob(job_id).status() or .result() when its state or result is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed or had already been committed; otherwise the app starts the later analysis turn for an unread final result. See the remote-compute-ssh skill for details.

Inside score_evo2.py, point HF_HOME at the provider's weight-cache mount (path is in compute_details) and set HF_HUB_OFFLINE=1 so the loader doesn't try to write refs/ into a read-only mount. Weight footprint: ~15 GB (7B), ~80 GB (40B).

Typical performance

Task 7B on H100 Notes
Model load (cached) ~5-7 min First call hydrates weights
score_sequences, 200×200bp ~10-20 s After load
generate, 1×512 nt ~15 s

Troubleshooting

Symptom Cause Fix
Transformer Engine not installed No FP8 — falls back to bf16 Informational only on non-H100; ignore
OOM on load 40B on <80 GB GPU Use evo2_7b or shard with device_map
HF tries to write refs/main HF_HOME points at RO mount Set HF_HUB_OFFLINE=1
dtype mismatch in score_sequences Passing tensors not strings Pass list[str]; the API tokenises for you

Next: pair with borzoi to predict track-level effects of the same variants.

Version History

  • 44394f0 Current 2026-09-11 11:10

Same Skill Collection

resources/skills/alphafold2/SKILL.md
resources/skills/boltz/SKILL.md
resources/skills/borzoi/SKILL.md
resources/skills/chai1/SKILL.md
resources/skills/compute-env-setup/SKILL.md
resources/skills/customize/SKILL.md
resources/skills/diffdock/SKILL.md
resources/skills/env-management/SKILL.md
resources/skills/fair-esm2/SKILL.md
resources/skills/figure-composer/SKILL.md
resources/skills/figure-style/SKILL.md
resources/skills/indication-dossier/SKILL.md
resources/skills/ligandmpnn/SKILL.md
resources/skills/literature-review/SKILL.md
resources/skills/openfold3/SKILL.md
resources/skills/paper-narrative/SKILL.md
resources/skills/proteinmpnn/SKILL.md
resources/skills/remote-compute-ssh/SKILL.md
resources/skills/scgpt/SKILL.md
resources/skills/scvi-tools/SKILL.md
resources/skills/self-awareness/SKILL.md
resources/skills/skill-creator/SKILL.md
resources/skills/solublempnn/SKILL.md
resources/skills/esmfold2/SKILL.md

Metadata

Files
0
Version
44394f0
Hash
52cf901c
Indexed
2026-09-11 11:10

ホーム - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-12 07:25
浙ICP备14020137号-1 $お客様$