Agent Skills › archestra-ai/archestra › archestra-dev-testing

archestra-dev-testing

GitHub

指导测试策略,判断测试层级与存在必要性。涵盖后端Vitest、E2E及前端规范,强调通过本地验证确保合并队列质量,优化CI成本。

.agents/skills/archestra-dev-testing/SKILL.md archestra-ai/archestra

Trigger Scenarios

决定测试该放在哪个层级 评估新测试的必要性 配置或优化CI/CD测试流程

Install

npx skills add archestra-ai/archestra --skill archestra-dev-testing -g -y
More Options

Non-standard path

npx skills add https://github.com/archestra-ai/archestra/tree/main/.agents/skills/archestra-dev-testing -g -y

Use without installing

npx skills use archestra-ai/archestra@archestra-dev-testing

指定 Agent (Claude Code)

npx skills add archestra-ai/archestra --skill archestra-dev-testing -a claude-code -g -y

安装 repo 全部 skill

npx skills add archestra-ai/archestra --all -g -y

预览 repo 内 skill

npx skills add archestra-ai/archestra --list

SKILL.md

Frontmatter
{
    "name": "archestra-dev-testing",
    "description": "Use for test selection and quality across backend, frontend, and e2e; load its backend reference for Vitest projects, mocking, DB fixtures, and performance."
}

What to test, and at which level

Run commands from platform/ unless specifically instructed otherwise.

This skill answers should this test exist, and where does it belong. Once you know the level, the mechanics live elsewhere:

  • For backend Vitest tests, read references/backend-tests.md for mocking rules, project selection, DB fixtures, and performance.
  • archestra-dev-e2e — Playwright fixtures, WireMock, selectors.
  • archestra-dev-frontend — component/query-hook conventions.

Frontend Vitest defaults to jsdom. For a test of pure logic whose runtime imports do not need browser APIs, put // @vitest-environment node at the top of the file. Check its dependency graph and run the file under Node before adding the directive; a .test.ts suffix alone does not establish that it is browser-free. This opt-in uses a shared worker module cache, so do not use vi.mock or leak mutable module/global state across files; the config rejects module mocks in this project. Keep UI, storage, canvas, and browser binary API tests on jsdom.

The one rule

A test earns its place by being able to fail for a reason you'd want to hear about.

CI time is a real budget. Every test runs on merge-queue attempts, and every test is code someone has to keep working during unrelated refactors. A test that can only fail when someone edits the literal it is compared against costs that budget and returns nothing.

Before writing a test, answer: what plausible mistake does this catch? If the answer is "someone deliberately changing this exact line", don't write it. If you cannot name the bug, there is no test to write.

Local checks and merge queue

Ordinary PR pushes run the PR policy checks; the expensive build, lint, test, and E2E jobs run on merge_group. Their skipped PR status is a queue-entry signal, not evidence that the code passed those checks. Before marking a PR ready, run the relevant checks locally from platform/ with pnpm (including focused tests for changed behavior) and review the PR policy results. Use the run-e2e label when browser coverage is worth getting before queue entry.

When the PR is ready, add it to the merge queue. If a queue check fails, read the failure, reproduce it locally, fix it, run the affected local checks, push the fix, and requeue. Keep the full failure picture from the queue run; do not blindly rerun until it passes. Release-please PRs are an exception: their platform lint/build job also runs before queue entry so generated release files can be committed.

Pick the level

Prefer the cheapest level that can actually catch the bug — but do not push a test down a level if that means mocking away the thing that would break.

Route-level integration tests are the default (backend)

This is the level to favour. A test that drives a real Fastify route through app.inject against the PGlite database exercises the schema, the auth middleware, endpoint permissions, the model layer, and the audit record all at once — the layers where bugs actually live.

const response = await app.inject({
  method: "POST",
  url: "/api/knowledge-bases",
  payload: { name: "Test KB" },
});
expect(response.statusCode).toBe(200);
expect(response.json().name).toBe("Test KB");

That assertion looks trivial in isolation, but it is not a fluff test: the value made a round trip through validation, authorization, and persistence. Asserting what came back out of a system is different from asserting what a literal says.

Never mock the database. Use the @/test fixtures (makeUser, makeOrganization, makeAgent, …) and real model calls.

MSW-backed integration tests are the default (frontend)

The frontend equivalent: render the real component tree with a real QueryClient and stub only the HTTP boundary with MSW. This covers the query hooks, loading and error states, cache invalidation and the rendered result together.

const server = setupServer(
  http.get(`${API_ORIGIN}/api/apps/:id`, () => HttpResponse.json(app)),
);

See frontend/src/app/a/[appId]/page.client.test.tsx for the established shape.

Unit tests, for real logic

Good unit tests cover branching, parsing, ordering, arithmetic, and edge cases — logic you can get wrong. shouldShowStickyBoundaryIndicator in frontend/src/components/chat/message-boundary-divider.test.tsx is a good one: a pure predicate with a genuine off-by-one boundary.

Reach for a unit test when the function has decisions in it. Not when it has none.

E2E tests, sparingly

E2E is the most expensive level: slow, flaky-prone, and delicate to maintain. Use E2E for a happy path that no cheaper level can cover — a real browser against a real stack, or real Kubernetes behaviour. Include failure/denial cases when only real infrastructure can establish the contract, such as NetworkPolicy enforcement.

Use route-level or MSW-backed integration tests for error branches, permission matrices, validation messages, and field-level behaviour when those levels preserve the boundary under test.

When a bug fix needs pinning, ask whether an integration test would catch the same regression. It usually would.

The fluff test anti-pattern

The trap, as it usually arrives: a change is made, a test "should" accompany it, and the fastest test to write is one that restates the code. It passes immediately, looks like diligence, and can never fail usefully.

The canonical shape — the assertion is the source, re-typed:

// Never. This asserts that JSX assigns props.
const comp = <MyComponent thing={false} />;
expect(comp.props.thing).toBe(false);

Concrete instances removed from this repo

Each of these shipped, ran on every CI build, and caught nothing:

A literal field compared to its own literal. Every knowledge-base connector declares its own type = "notion" as const, and eight tests asserted exactly that back:

// Removed — cannot fail unless someone edits the line above it.
it("has the correct type", () => {
  expect(new NotionConnector().type).toBe("notion");
});

BaseConnector types that field as ConnectorType and registry.ts keys by it, so a rename is a compile error or a registry break — both louder than a test.

But check the claim before you lean on it. The same suites also asserted expect(connector.supportsPermissionSync).toBe(true), which looks like the identical pattern and is not. That field is a plain boolean: flipping it to false compiles cleanly, and — verified by doing it — every one of the connector's own tests still passes. Meanwhile five call sites gate real behaviour on it, so the flip would silently stop permission syncing for that source.

The fix was not to keep fifteen copies of the literal, and not to delete the coverage either — it was to move it to the level where it has a consumer:

// One capability matrix, pinned where the scheduler actually reads it.
test("pins the connector types that implement permission sync", () => {
  expect(getPermissionSyncConnectorTypes().filter((t) => t !== "perforce").sort())
    .toEqual([...]);
});

perforce is excluded because it computes the flag from isK8sConfigured(). The lesson generalises: "the type system already covers this" and "the neighbouring tests already cover this" are both claims you can test in about two minutes by breaking the code and running the suite. Do that before deleting anything.

Prop pass-through. Asserting that a className prop reaches the element it is spread onto tests JSX, not the component:

// Removed — a property of the framework, not of this component.
it("applies custom className to combobox", () => {
  render(<SearchableMultiSelect className="my-custom-class" {...rest} />);
  expect(screen.getByRole("combobox")).toHaveClass("my-custom-class");
});

Styling copied from the source. Asserting a component's own utility classes turns every restyle into a test edit while catching no user-visible bug:

// Removed — the test is a copy of the className string in the component.
expect(status).toHaveClass("min-h-[calc(var(--visual-viewport-height,100dvh)-12rem)]");
expect(wrapper).toHaveClass("sm:w-[320px]");

Assertions the type system already proves. Not a removal, but the pattern to watch for: resourceLabels is a Record<Resource, string>, so a bare "every resource has a label" existence check cannot fail at runtime. Where a type guarantees the shape, a test earns its place only by asserting what the type cannot — an empty string, say, which is why shared/permission.types.test.ts also checks length and category membership and stays.

Class names are not automatically fluff

The distinction is whether the class carries a contract or is just styling:

  • Keep: motion-reduce:animate-none (an accessibility escape hatch), secret-masked (a secret is actually masked), a relative wrapper that a regression once broke by dropping (see search-input.test.tsx).
  • Drop: size-8, text-sm, sm:w-[320px] — sizing and colour with no behaviour attached.

Prefer asserting the user-visible fact instead: role, accessible name, visible text, enabled/disabled, what is on screen and what is not.

Never leave a disabled test behind

A permanently skipped test is worse than no test: it reads as coverage, runs never, and rots against the code it claims to describe. This repo had a describe.skip("AgentForm") with no reason or link; when un-skipped, two of its three tests passed — real coverage that had been off for months — and the third no longer matched the component at all.

  • Fix it, or delete it. If it must stay off, say why and link the tracking issue (oauth-self-hosted.spec.ts does this correctly).
  • A test.skip("...", () => {}) with an empty body is a comment. Write a comment.
  • Conditional skips for genuine environment gates (test.skip(!byosEnabled, "…")) are fine — they are runtime guards, not dead code.

Duplication: cheap and load-bearing vs. expensive and redundant

Repetition across files is not automatically waste. Weigh what it costs against what it covers:

  • Keep the twelve "rejects an empty batch" bulk-route tests. They all exercise one shared BulkIdsSchema, but each also pins that its own route wired that schema up, and each is one extra app.inject on an app the suite already built — microseconds.
  • Keep the parallel team-scope suites for mcp-oauth-clients and llm-oauth-clients. Eleven of their test bodies are byte-identical, but they guard two independent authorization surfaces with separately-implemented scope sequencing; deleting either half would let a privilege-escalation bug through on one of them.
  • Remove repetition that re-asserts a shared literal per instantiation with no per-instance wiring to prove — the connector flags above.

The question is never "is this duplicated" but "does the second copy have its own way to fail".

Reuse setup without hiding the behavior

Before copying setup into a new test, check the existing fixtures and data factories. Backend database tests should use @/test fixtures rather than repeating table inserts; add a focused fixture when the same meaningful setup recurs. Keep database-free *.unit.test.ts files free of @/test and database imports. Frontend integration tests should reuse frontend/tests-integration/fixtures.ts and the factories in frontend/src/mocks/data/ before adding page, MSW, or entity setup of their own. Batch independent MSW overrides with mswControl.registerMany when a scenario needs several of them. Keep each scenario's requests and assertions visible in its test so the behavior remains clear.

Checklist before adding a test

  1. Name the bug it catches. Can't? Don't write it.
  2. Could the type system, a schema, or the compiler already catch it? Then skip it, or narrow the test to the part they can't prove.
  3. Pick the cheapest level that still exercises the risky part — but don't mock away the thing under test to get there.
  4. Assert observable behaviour: responses, rendered output, persisted rows, emitted events. Not internal shape, not styling, not that a prop arrived.
  5. E2E for happy paths or infrastructure contracts nothing cheaper can reach.
  6. If it can't be made to pass, don't commit it skipped.

Version History

  • 3bed5da Current 2026-09-27 12:11

    将昂贵验证移至合并队列以节省CI资源;调整测试分片与构建时间;更新后端测试参考链接。

  • 6290aab 2026-09-11 16:19

Same Skill Collection

.agents/skills/archestra-dev-backend-tests/SKILL.md
.agents/skills/archestra-dev-backend/SKILL.md
.agents/skills/archestra-dev-bench-analysis/SKILL.md
.agents/skills/archestra-dev-e2e/SKILL.md
.agents/skills/archestra-dev-frontend/SKILL.md
.agents/skills/archestra-dev-investigate/SKILL.md
.agents/skills/archestra-dev-llm-providers/SKILL.md
.agents/skills/archestra-dev-migrations/SKILL.md
.agents/skills/archestra-dev-observability/SKILL.md
.agents/skills/archestra-dev-override-sweep/SKILL.md
.agents/skills/archestra-dev-resource-cleanup/SKILL.md
.agents/skills/archestra-dev-rust-napi/SKILL.md
.agents/skills/archestra-docs-writer/SKILL.md
.agents/skills/archestra-mcp-catalog-entry/SKILL.md
.agents/skills/archestra-pr-hygiene/SKILL.md
.agents/skills/managing-archestra-releases/SKILL.md
.claude/skills/archestra-dev-backend-tests/SKILL.md
.claude/skills/archestra-dev-backend/SKILL.md
.claude/skills/archestra-dev-bench-analysis/SKILL.md
.claude/skills/archestra-dev-e2e/SKILL.md
.claude/skills/archestra-dev-frontend/SKILL.md
.claude/skills/archestra-dev-investigate/SKILL.md
.claude/skills/archestra-dev-llm-providers/SKILL.md
.claude/skills/archestra-dev-migrations/SKILL.md
.claude/skills/archestra-dev-observability/SKILL.md
.claude/skills/archestra-dev-override-sweep/SKILL.md
.claude/skills/archestra-dev-rust-napi/SKILL.md
.claude/skills/archestra-dev-testing/SKILL.md
.claude/skills/archestra-docs-writer/SKILL.md
.claude/skills/archestra-mcp-catalog-entry/SKILL.md
.codex/skills/archestra-dev-backend-tests/SKILL.md
.codex/skills/archestra-dev-backend/SKILL.md
.codex/skills/archestra-dev-bench-analysis/SKILL.md
.codex/skills/archestra-dev-e2e/SKILL.md
.codex/skills/archestra-dev-frontend/SKILL.md
.codex/skills/archestra-dev-investigate/SKILL.md
.codex/skills/archestra-dev-llm-providers/SKILL.md
.codex/skills/archestra-dev-migrations/SKILL.md
.codex/skills/archestra-dev-observability/SKILL.md
.codex/skills/archestra-dev-override-sweep/SKILL.md
.codex/skills/archestra-dev-rust-napi/SKILL.md
.codex/skills/archestra-dev-testing/SKILL.md
.codex/skills/archestra-docs-writer/SKILL.md
.codex/skills/archestra-mcp-catalog-entry/SKILL.md
.codex/skills/archestra-pr-hygiene/SKILL.md
ai-labs/skills/access-request-intake/SKILL.md
ai-labs/skills/cipher-decoder/SKILL.md
ai-labs/skills/sales-ledger/SKILL.md
migration-kit/SKILL.md

Metadata

Files
0
Version
3bed5da
Hash
6a4712f3
Indexed
2026-09-11 16:19

Home - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-28 15:59
浙ICP备14020137号-1