mime-detection-routing
GitHub提供MIME类型检测、提取器路由及格式注册表的开发指南。涵盖路径与字节检测逻辑、优先级选择、通配符支持,以及新增格式的标准化流程,用于调试路由错误或扩展系统能力。
Trigger Scenarios
Install
npx skills add xberg-io/xberg --skill mime-detection-routing -g -y
SKILL.md
Frontmatter
{
"name": "mime-detection-routing",
"description": "MIME type detection and extractor routing in core\/mime.rs — the FORMATS registry that EXT_TO_MIME and SUPPORTED_MIME_TYPES are derived from, the path-based and bytes-based detection functions, priority-based registry selection, wildcard MIME families, and the real procedure for adding a format. Load when adding a format, wiring an extractor to a MIME type, or debugging why a file routes to the wrong (or no) extractor."
}
MIME Detection & Routing
Detection Flow
Policy -> content and/or extension evidence -> validate_mime_type -> registry.get(mime) -> extractor
Key Functions
| Function | Location | Behaviour |
|---|---|---|
detect_mime_type(path, check_exists: bool) |
core/mime.rs |
Path-based only — never reads bytes. Lowercased extension → EXT_TO_MIME, then tree-sitter extension detection (feature tree-sitter), then mime_guess::from_path. check_exists gates a file-existence check, not content inspection. |
detect_mime_type_from_bytes(bytes) |
core/mime.rs |
Magic-number detection via the infer crate. The only content-sniffing entry point. |
validate_mime_type(mime) |
core/mime.rs |
Parses the media type, matches its case-insensitive essence against SUPPORTED_MIME_TYPES, and returns the registered MIME spelling. Parameters such as charset do not affect extractor routing. It does not consult the extractor registry. |
The FORMATS registry is the single source of truth
FORMATS: &[FormatEntry { extensions, mime_type, aliases }] in core/mime.rs.
EXT_TO_MIME and SUPPORTED_MIME_TYPES are LazyLocks derived from it by iteration —
there is no m.insert call site to add to, and hand-editing either is impossible.
The full registry publishes 106 formats, 140 unique extensions, and 53 aliases, verified by
scripts/sync_supported_counts.py verify. The published count constants describe that static
registry; runtime availability is its intersection with registered extractors. Extension lookup is
case-insensitive (the extension is lowercased before the map hit).
Registry Selection
let registry = get_document_extractor_registry(); // plugins/registry/mod.rs
let guard = registry.read()?;
let extractor: Arc<dyn DocumentExtractor> = guard.get(mime_type)?; // Result, not Option
DocumentExtractorRegistry::get (plugins/registry/extractor.rs) returns the
highest-priority() extractor for the MIME type, and returns Err — not None — when none
matches.
Wildcard Support
An extractor may register a family: "image/*" matches image/png, image/jpeg, and so on
(prefix match on a registered type ending in /*).
Adding a New Format
- Add one
FormatEntrytoFORMATSincrates/xberg/src/core/mime.rs.EXT_TO_MIMEandSUPPORTED_MIME_TYPESupdate automatically. - Run
scripts/sync_supported_counts.py syncto update published count claims, then run itsverifycommand. - Implement
InternalDocumentExtractor(notDocumentExtractor— seeplugin-architecture-patterns) withsupported_mime_types()returning the MIME. - Register in
crates/xberg/src/extractors/mod.rs::register_default_extractors().
Critical Rules
- Call
validate_mime_type()before extraction — but do not treat it as proof an extractor exists. - Extension lookup is case-insensitive.
detect_mime_typeinspects no content. Extraction defaults toPreferContent, which performs bounded content inspection and falls back to a supported extension. UseContentOnlywhen the filename must be ignored; useTrustExtensiononly for trusted sources.- A specific explicit MIME type is authoritative.
application/octet-streamis the exception: it is a generic placeholder and triggers policy-based detection. - Never edit
EXT_TO_MIMEorSUPPORTED_MIME_TYPES— editFORMATS.
Version History
-
d8e4815
Current 2026-08-28 18:30
重构为基于FORMATS单一事实源的架构,移除手动编辑映射,更新检测函数行为描述及新增格式步骤。
- 531e0f7 2026-08-20 07:47


