Agent Skills › cbrock84/headcount › assessment-design

assessment-design

GitHub

专注于教育评估设计,指导如何构建有效的测验题目与评估体系。涵盖明确测量目标、匹配题型与能力、编写诊断性干扰项及基于多题证据判断掌握度,旨在提升评估的准确性与教学价值。

verticals/education/skills/education/assessment-design/SKILL.md cbrock84/headcount

Trigger Scenarios

编写或审查测验题目 设计单元检测或诊断性评估 分析考试成绩与课堂表现不符的原因

Install

npx skills add cbrock84/headcount --skill assessment-design -g -y
More Options

Non-standard path

npx skills add https://github.com/cbrock84/headcount/tree/main/verticals/education/skills/education/assessment-design -g -y

Use without installing

npx skills use cbrock84/headcount@assessment-design

指定 Agent (Claude Code)

npx skills add cbrock84/headcount --skill assessment-design -a claude-code -g -y

安装 repo 全部 skill

npx skills add cbrock84/headcount --all -g -y

预览 repo 内 skill

npx skills add cbrock84/headcount --list

SKILL.md

Frontmatter
{
    "name": "assessment-design",
    "description": "Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from one that measures reading speed or test-taking. Use this to write or review assessment items, design a unit check or diagnostic, work out why scores do not match classroom performance, or decide how much evidence a claim about mastery actually needs."
}

Assessment design

Every question measures something. The work is making sure it measures the thing you meant, because the alternatives — reading speed, familiarity with the format, willingness to guess — are always available and are usually easier for the student to use.

Name the inference before writing the item

An assessment is an argument: the student did this, therefore they can do that. The argument is where assessments fail, and it fails silently.

Write down the claim first — "can decompose a two-digit number into tens and ones" — then ask what performance would be evidence for it, and what performance would be evidence against. An item that a student who lacks the skill can still get right is not evidence. An item that a student who has the skill can still get wrong, for reasons unrelated to it, is worse: it produces a false negative that gets acted on.

The most common unrelated reason is reading. Any item whose stem is harder to read than the skill is to perform has quietly become a reading assessment.

Match the item type to the claim

  • Selected response (multiple choice, matching, true/false) is efficient and can only ever provide evidence of recognition. A student who can recognize the correct answer cannot be assumed to produce it.
  • Constructed response shows the path, which is what makes partial understanding visible. It costs scoring time and needs a rubric written before the responses arrive, not after.
  • Performance tasks are the only honest evidence for anything described as applying, investigating or designing — most NGSS performance expectations and every C3 inquiry, for instance, cannot be assessed by selected response at all.

Mismatch is the usual failure: a standard describing explanation assessed by a multiple-choice item, which measures whether the student can pick an explanation someone else wrote.

Distractors are the diagnostic

In a well-built multiple-choice item, each wrong answer is the result of a specific, predictable error. Then the pattern of wrong answers says what to reteach, and the item earns its place.

Distractors that are merely wrong — a random number, an obviously absurd option — turn a four-option item into a two-option one and tell you nothing beyond right or wrong.

The mechanical tells of a weak item, all of which students learn to exploit long before they learn the content:

  • The longest or most qualified option is correct.
  • One option is grammatically inconsistent with the stem.
  • "All of the above" appears, and is usually correct.
  • Two options are synonyms, so neither can be right.
  • The correct answer repeats wording from the stem.

One item is not evidence

A single item carries noise — a misread word, a slip, a lucky guess — that swamps the signal for any individual student. Inferring mastery from one response is the most common measurement error in classroom material, and the most consequential, because it gets recorded.

Several items per claim, varied in surface form so that recognition of the format is not what is being measured. If a claim is worth recording against a student's name, it is worth three items.

Spacing matters as much as quantity: performance on a skill immediately after it is taught measures something closer to short-term recall than to learning. The same item a fortnight later measures more.

Formative and summative are different products

Formative assessment exists to change what happens next, which means it must be quick, frequent, low-stakes, and read immediately. An assessment that takes a week to score cannot be formative whatever it is called.

Summative assessment exists to record a judgment, which means it must be defensible: enough items, a rubric written in advance, and conditions that were the same for everyone.

Material sold as practice is usually formative in function and summative in appearance, which is worth stating plainly to a buyer. education:learning-materials-design covers the practice sequence this sits inside.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Never

  • Write an item before naming the claim it is evidence for.
  • Assess an explanation standard with a recognition item.
  • Build distractors that are wrong without being diagnostic.
  • Record mastery from a single response.
  • Make the stem harder to read than the skill is to perform.

Version History

  • 98d1c17 Current 2026-09-22 00:34

Same Skill Collection

plugins/corporate-strategy/skills/chief-strategy-officer/SKILL.md
plugins/corporate-strategy/skills/market-entry/SKILL.md
plugins/corporate-strategy/skills/mergers-and-acquisitions/SKILL.md
plugins/corporate-strategy/skills/portfolio-strategy/SKILL.md
plugins/corporate-strategy/skills/scenario-planning/SKILL.md
plugins/corporate-strategy/skills/strategic-alliances/SKILL.md
plugins/customer-experience/skills/chief-customer-officer/SKILL.md
plugins/customer-experience/skills/customer-onboarding-and-implementation/SKILL.md
plugins/customer-experience/skills/customer-success-management/SKILL.md
plugins/customer-experience/skills/escalation-management/SKILL.md
plugins/customer-experience/skills/self-service-and-knowledge/SKILL.md
plugins/customer-experience/skills/support-operations/SKILL.md
plugins/customer-experience/skills/voice-of-customer/SKILL.md
plugins/data-analytics/skills/ai-ml-governance/SKILL.md
plugins/data-analytics/skills/business-intelligence/SKILL.md
plugins/data-analytics/skills/chief-data-officer/SKILL.md
plugins/data-analytics/skills/data-engineering/SKILL.md
plugins/data-analytics/skills/data-governance/SKILL.md
plugins/data-analytics/skills/data-modeling/SKILL.md
plugins/demand-generation/skills/ai-search-optimization/SKILL.md
plugins/demand-generation/skills/app-store-optimization/SKILL.md
plugins/demand-generation/skills/experimentation/SKILL.md
plugins/demand-generation/skills/landing-page-cro-expert/SKILL.md
plugins/demand-generation/skills/lead-capture/SKILL.md
plugins/demand-generation/skills/lifecycle-messaging/SKILL.md
plugins/demand-generation/skills/listing-distribution/SKILL.md
plugins/demand-generation/skills/marketing-analytics/SKILL.md
plugins/demand-generation/skills/paid-advertising/SKILL.md
plugins/demand-generation/skills/programmatic-seo/SKILL.md
plugins/demand-generation/skills/seo-strategy/SKILL.md
plugins/executive/skills/ai-research-analyst/SKILL.md
plugins/executive/skills/business-growth-consultant/SKILL.md
plugins/executive/skills/chief-executive/SKILL.md
plugins/executive/skills/fundraising-and-investor-relations/SKILL.md
plugins/executive/skills/saas-idea-validator/SKILL.md
plugins/finance/skills/budgeting-and-forecasting/SKILL.md
plugins/finance/skills/capital-allocation/SKILL.md
plugins/finance/skills/capital-structure-and-covenants/SKILL.md
plugins/finance/skills/cost-accounting/SKILL.md
plugins/finance/skills/financial-modeling/SKILL.md
plugins/finance/skills/financial-reporting-and-close/SKILL.md
plugins/finance/skills/financial-statement-analysis/SKILL.md
plugins/finance/skills/internal-controls-and-audit/SKILL.md
plugins/finance/skills/revenue-recognition/SKILL.md
plugins/finance/skills/tax/SKILL.md
plugins/finance/skills/treasury-and-liquidity/SKILL.md
plugins/finance/skills/unit-economics/SKILL.md
plugins/it-operations/skills/backup-and-recovery/SKILL.md
plugins/it-operations/skills/chief-information-officer/SKILL.md
plugins/it-operations/skills/cloud-administration/SKILL.md

Metadata

Files
0
Version
98d1c17
Hash
5b9e360b
Indexed
2026-09-22 00:34

Home - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-28 05:36
浙ICP备14020137号-1