Agent Skillstutti-os/tutti › tutti-record-agent-session-replay

tutti-record-agent-session-replay

GitHub

用于Qualify Tutti Agent Session Replay Cassettes。基于脚本驱动而非UI录制,执行Live Record、结构审计、Fresh Replay以验证Cassette有效性,并支持新Provider捕获及缺陷诊断。

.codex/skills/tutti-record-agent-session-replay/SKILL.md tutti-os/tutti

触发场景

Session Replay cassette qualification CDP-executed scenario replay diagnosis New Provider capture support

安装

npx skills add tutti-os/tutti --skill tutti-record-agent-session-replay -g -y
更多选项

非标准路径

npx skills add https://github.com/tutti-os/tutti/tree/main/.codex/skills/tutti-record-agent-session-replay -g -y

不安装直接使用

npx skills use tutti-os/tutti@tutti-record-agent-session-replay

指定 Agent (Claude Code)

npx skills add tutti-os/tutti --skill tutti-record-agent-session-replay -a claude-code -g -y

安装 repo 全部 skill

npx skills add tutti-os/tutti --all -g -y

预览 repo 内 skill

npx skills add tutti-os/tutti --list

SKILL.md

Frontmatter
{
    "name": "tutti-record-agent-session-replay",
    "description": "From a Tutti checkout, run, audit, freshly replay, publish, or diagnose Session Replay cassettes that are driven by case-repository scenario scripts (CDP), not by interactive UI recording. Use for real-Provider capture while a scenario.mjs executes, cassette transport or semantic-state mismatches, fresh replay qualification, AgentGUI replay evidence, and product-side support for a new Agent Target beyond local:codex \/ local:claude-code. Do not use to author Case metadata or scenario scripts (case repository write-replay-case), or for executionKind \"ui\" Cases."
}

Qualify Tutti Agent Session Replay Cassettes

Work from the Tutti checkout. Keep product implementation and the generic runner in Tutti; keep Case metadata, scenario scripts, fixtures, qualified Cassettes, and evidence in the external case repository (tutti-os/tutti-replay).

Mental model (script-first, not UI recording):

  1. Humans/agents write a deterministic scenarios/*.mjs (prepare / drive / assert) in the case repository — that script is the recording plan.
  2. Record means: Tutti runner launches Desktop, CDP-executes that script against a live Provider, and captures the Cassette. There is no separate click-to-record UI workflow for Session Replay.
  3. Day-to-day Record/Replay is usually triggered from the case repository QA console; this skill is for Tutti-side CLI qualification, diagnosis, runner or Replay product defects, and new Provider capture support.

Prove qualification in this order:

existing scenario script -> live Record (script + Provider) -> structural audit -> fresh Replay -> optional publication

Never call a Cassette qualified until Record, audit, and a fresh isolated Replay have all passed. If the scenario script itself is missing or wrong, stop and use the case repository write-replay-case skill — do not invent Cases inside Tutti.

Start the QA console (case repository)

Browsing Cases, Test Plans, and one-click Record/Replay (which run the same scenario scripts) live in the case repository (sibling checkout, commonly ../tutti-replay; GitHub: tutti-os/tutti-replay).

From the case repository root:

pnpm install
pnpm dev

Open only http://127.0.0.1:3333 (never the API port :3334). In the UI, set the Tutti checkout absolute path, create a Test Plan, then Record or Replay. First-time machine setup: that repository's SETUP.md. Authoring or mirroring scenario scripts: .agents/skills/write-replay-case/.

For a long-lived LAN service on macOS use pnpm replay:service install (port 2333); do not run pnpm dev and the stable service at the same time.

Establish scope

  1. Read the Tutti root and closest AGENTS.md files.
  2. Read docs/architecture/agent-session-replay.md (Provider support: developer recording currently accepts local:codex and local:claude-code only).
  3. For AgentGUI behavior, also read docs/architecture/agent-gui-node.md and packages/agent/gui/AGENTS.md.
  4. For Session, Turn, Goal, or runtime-operation lifecycle behavior, read packages/agent/host/README.md; lifecycle semantics remain in Host.
  5. Inspect git status --short in both repositories and preserve pre-existing work.
  6. Resolve the case repository from the user-provided path, the configured cases path, or the sibling ../tutti-replay checkout. Do not guess another location if none exists.

Read the selected Case before planning:

  • cases/<case-id>/case.json
  • every cases/<case-id>/scenarios/*.mjs except *.impl.mjs
  • referenced shared scenario helpers and runtime-fixtures/
  • existing cassettes/, evidence/, and relevant Run artifacts
  • the case repository's README.md and CONTEXT.md when publication or Case lifecycle is involved

If case.json declares executionKind: "ui", stop this Session Replay workflow. Pure UI Cases are still script-driven (defineUiScenario + CDP), but they use ui-drive and publish ui/ screenshots — they do not Record Provider Cassettes. Author and run them via the case repository write-replay-case skill and the QA console; do not use this skill's --record / Cassette audit path for them.

Use CDP through Tutti's repository runner. Do not use Computer Use unless the user explicitly requests it.

Maintain the ownership boundary

  • Add or update Case scenarios only under cases/<case-id>/scenarios/*.mjs in the case repository.
  • Put reusable scenario helpers in that repository's scenario-runtime/.
  • Do not add Case registries, Case-specific scenarios, fixtures, or qualified Cassettes to Tutti.
  • Change Tutti only for generic product, runner, protocol, or Replay defects.
  • Fix root causes. Do not relax transport matching, semantic verification, terminal assertions, or checkpoint requirements to accept a broken Case.

Add a new Provider (product vs case repository)

Today Replay recording targets are local:codex and local:claude-code. The shared Session Replay core is provider-neutral, but each new Agent Target still needs Tutti capture + fail-closed playback before any Case work is useful.

Split work explicitly:

Layer Where Who Scope
Product Tutti experienced / mentored Adapter capture, projected tape, portability, structural audit, outbound verification, input-unit barriers, isolated Provider home, deterministic fail-closed Replay for local:<provider>
Cases case repository can hand to intern after product gate providerProfiles, KNOWN_PROVIDERS, defineMirroredRecordScenario mirrors, Record via console, publish cassettes

Do not start by writing Cases for an unsupported Provider. Product must accept --agent-target-id local:<provider> for Record and Replay first.

Tutti checklist (this repository)

Copy and tick:

- [ ] 1. Provider adapter can Record real traffic into a Cassette
- [ ] 2. Projected tape + portability (paths/homes) match session-replay contract
- [ ] 3. Structural audit passes (manifest, frames, activity causality)
- [ ] 4. Fresh isolated Replay is fail-closed (no live Provider fallback)
- [ ] 5. Runner accepts --agent-target-id local:<provider>
- [ ] 6. docs/architecture/agent-session-replay.md Provider support updated
- [ ] 7. Hand off to case repository write-replay-case for mirrors + console Record

Prove with one smoke scenario from the case repository (or a temporary scenario file) using the Tutti runner Record → audit → fresh Replay loop below. Keep account secrets out of logs and reports.

Case repository handoff

After the product gate is green, follow that repository's write-replay-case skill section on multi-provider mirrors. Typical touch points there (not in Tutti): scenario-runtime/shared.mjs providerProfiles, src/shared/agent-target.ts KNOWN_PROVIDERS, mirrored *.impl.mjs + variant Case dirs, then console Record/Replay publication.

Scenario scripts (owned by the case repository)

Session Replay drive logic is authored as scripts, not captured from manual UI interaction. Prefer the case repository's write-replay-case skill and defineMirroredRecordScenario / defineRecordScenario helpers.

When diagnosing or qualifying from Tutti, still require that the loaded scenario:

  • exports prepare, drive, and assert;
  • sets Provider via profile/helper (not a hard-coded Codex-only identity);
  • sets every behavior-affecting composer default explicitly;
  • uses one stable prompt per intended Turn, with exact markers / final tokens;
  • uses accessible labels, test IDs, or semantic DOM state instead of coordinates;
  • waits before each interaction; answers each approval/question/plan once;
  • asserts a terminal state with no enabled stale controls;
  • declares expectedRecordingMode when continuing an existing Session.

For question cards, the script must trigger the Provider's real user-input request. For plan Cases, wait for the completed plan and implementation decision before driving the real action.

If the script needs to change, edit it in the case repository — then re-Record.

Record = execute the scenario script with a live Provider

From the Tutti root, run the repository runner so it CDP-drives the scenario file and captures a Cassette. Derive scenario ID, Cassette name, and Agent Target from the loaded scenario:

pnpm e2e:agent-gui -- \
  --record .tmp/cassettes/<cassette-name> \
  --scenario <scenario-id> \
  --scenario-file <case-repository>/cases/<case-id>/scenarios/<scenario-id>.mjs \
  --agent-target-id <agent-target-id> \
  --keep-runtime \
  --timeout-ms 300000 \
  --stall-timeout-ms 60000

Prefer the case console Test Plan「录制」for routine work; use this CLI when debugging runner/Replay behavior or when the console is unavailable.

Omit --headless while debugging; add it for unattended execution. Keep the runtime only long enough to inspect or collect its artifacts.

Inspect failures from the smallest relevant evidence set:

  • record screenshots and checkpoint screenshots;
  • logs/desktop.log;
  • state/logs/tuttid.log;
  • state/tuttid.db through targeted queries;
  • the incomplete Cassette and a small decoded Provider-frame window.

Do not dump an entire Provider stream or expose account data.

Audit each Cassette

Run the bundled structural audit from the Tutti root:

node .codex/skills/tutti-record-agent-session-replay/scripts/audit-cassette.mjs \
  .tmp/cassettes/<cassette-name>

Require:

  • Cassette inventory, hashes, and size policy verify;
  • the Provider manifest is complete;
  • global and per-connection frame sequences are continuous;
  • Activity sequences are continuous;
  • intent-to-effect causality satisfies packages/agent/session-replay/activity-contract.json;
  • expected interactions, plan decisions, tools, exits, terminal Turns, and final response state match case.json and the scenario assertions.

The bundled audit proves structural invariants and emits a semantic summary. It does not replace Case-specific assertions or a fresh Replay.

Run a fresh Replay

Replay from a fresh isolated Tutti runtime and pass the scenario so screenshot settling and terminal assertions run:

pnpm e2e:agent-gui -- \
  --replay .tmp/cassettes/<cassette-name> \
  --scenario <scenario-id> \
  --scenario-file <case-repository>/cases/<case-id>/scenarios/<scenario-id>.mjs \
  --screenshot-checkpoints \
  --keep-runtime \
  --timeout-ms 300000 \
  --stall-timeout-ms 60000

Require the runner's replay passed result, all planned checkpoints, the expected AgentGUI terminal state, and a fully drained Provider transport. Provider transport remains fail-closed; only repository-declared observer-only probes may yield to causal traffic.

When one Case owns multiple scenarios, Record and audit every resulting Cassette, then qualify them together through one Replay Workspace. Let the case repository workflow generate the workspace manifest; do not invent a second Case registry in Tutti.

Publish only qualified artifacts

Prefer the case repository's publication workflow because it records the Run, archives prior artifacts, and publishes cassettes/ plus evidence/ only after qualification.

If the user explicitly requests manual publication:

  1. Stage all new Cassettes and evidence without touching the qualified copies.
  2. Verify every staged Cassette completed Record, audit, and fresh Replay.
  3. Replace cases/<case-id>/cassettes/ and evidence/ as one Case operation, preserving the previous qualified artifacts in the case repository's archive convention.
  4. Re-read every published manifest and evidence directory.

Do not infer Case lifecycle from a successful command. Run status, manual acceptance, and case.json lifecycle status are distinct. Do not set status: "confirmed" or acceptedAt unless the case repository's acceptance requirements have been satisfied.

Finish

  • If Tutti implementation changed, run the validation selected by docs/conventions/testing.md plus any closest-area checks.
  • If case metadata or scenarios changed, run the case repository's metadata, scenario, and type checks.
  • Recheck both worktrees and separate pre-existing changes from this task.
  • Perform the Tutti documentation-impact check.

Report:

  • Case and Cassette names;
  • Record, audit, and fresh Replay results;
  • Provider-frame, Activity, interaction, tool, Turn, and final-state summaries;
  • Tutti implementation changes and case repository artifact changes;
  • changed-line distribution by functional area, excluding pre-existing work;
  • documentation impact;
  • failed gates and unimplemented scope.

版本历史

  • 0f78158 当前 2026-08-08 11:33

    将技能框架明确为脚本驱动的资格验证,区分产品端与案例仓库的工作边界,强调基于脚本+实时Provider捕获生成cassettes而非交互式UI录制。

  • b63c209 2026-08-02 23:45

同 Skill 集合

.codex/skills/tutti-app-release/SKILL.md
.codex/skills/tutti-architecture-review/SKILL.md
packages/ui/system/agent/tutti-ui-system/SKILL.md
services/tuttid/service/workspace/agent_workspace_app_reference/SKILL.md
.codex/skills/analyze-performance-traces/SKILL.md
services/tuttid/service/workspace/app_factory_reference/SKILL.md

元信息

文件数
0
版本
0f78158
Hash
9814ff26
收录时间
2026-08-02 23:45

首页 - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-09 16:35
浙ICP备14020137号-1 $访客地图$