jobseek-error-review
GitHub每日只读审查 Jobseek 爬虫错误,去重已知问题,分析 GitHub Issue,仅对新颖、回归或激增的错误创建或更新 Issue,并生成 Markdown 报告。
Trigger Scenarios
Install
npx skills add colophon-group/jobseek --skill jobseek-error-review -g -y
SKILL.md
Frontmatter
{
"name": "jobseek-error-review",
"description": "Daily read-only review of Jobseek crawler errors on Hetzner; dedupe known issues, inspect GitHub issues, collect evidence, and file or update issues only for novel, regressing, spiking, or incident classes."
}
Jobseek Error Review
Use this skill for the daily crawler error review. Operate from the repo state and these instructions; do not rely on prior conversation context.
Prefer running this through the Hetzner local Codex runner or a local Codex CLI session from the repository root. The Hetzner runner uses an isolated worktree and a root-collected redacted evidence bundle, which keeps routine output isolated from active development work while avoiding direct Docker or deploy-env access from the Codex process. This routine should use subscription-backed Codex execution where possible. Do not implement it as a GitHub Actions workflow.
The legacy Claude slash command at
.claude/commands/jobseek-error-review.md remains a compatibility fallback.
Keep behavior equivalent when using that fallback, but treat this skill as the
Codex-first runbook.
Do not spawn subagents by default. They cost extra tokens and are only useful
when the current run has large independent evidence sets and the user
explicitly asks for them. If subagents are used, keep them read-only and make
the main agent responsible for dedupe, classification, redaction, and all
GitHub write decisions. Use the project custom agent
jobseek-error-review-researcher (GPT-5.6 Terra, high reasoning) for each
bounded evidence set.
Mission
Scan the last 24 hours of errors on the crawler Hetzner box, classify each error class against prior reviews and prior GitHub issues, write a dated Markdown report, and file or update GitHub issues only for classes that meet the filing criteria.
Model issue quality after prior examples #2622, #2621, #2470, and #2431. Never mutate remote host state. GitHub writes are allowed only when the class is novel, regressing, spiking, or an incident after dedupe.
Target
Use the root-collected runner bundle by default. A local manual fallback may
resolve the production crawler host from approved operator configuration, but
must not copy host addresses into reports, traces, prompts, or GitHub content.
The Compose file reference is /home/deploy/docker-compose.yml.
Long-running services. Collect logs with explicit --since and --until:
deploy-worker-1-1- HTTP workerdeploy-worker-2-1- HTTP workerdeploy-worker-3-1- HTTP workerdeploy-browser-1-1- Playwright/Lightpanda workerdeploy-exporter-1- Postgres to Supabase and Typesense CDCdeploy-drain-1- R2 description uploaderdeploy-redis-1- local queue; only OOM andlevel=errorlinesdeploy-alloy-1- log/metric shipper; only Alloy's ownlevel=error
Ephemeral one-shots are exited containers created inside the review window
whose image matches ghcr.io/*/jobseek-crawler:*. Enumerate them with:
docker ps -a --filter status=exited --format '{{.Names}} {{.Image}} {{.CreatedAt}} {{.Status}}'
These include crawler refresh-typesense, backfill-typesense, sync,
notify-indexnow, and reconcile runs. Grab logs with docker logs <id>
before exited container logs age out. Ignore interactive debug containers
such as tesla-debug, stupefied_hofstadter, and goofy_haibt.
Inputs
-
Use a window of the last 24 hours ending at UTC now, rounded down to the minute. Use explicit
--sinceand--untilvalues so log collection exactly matches the report header. When the Hetzner runner provides a redacted evidence bundle, readmanifest.json,host/docker-inspect-state.txt, andhost/docker-cgroup-memory.jsonbefore classifying host incidents. Also readhost/docker-lifecycle.jsonl, the durable allowlisted stream ofcreate/start/restart/die/oom/kill/stop/destroyevents. These files record container identity, exit code or signal, event time, restart count, and cgroup-v2memory.eventscounters without exposing arbitrary Docker labels. Readhost/maintenance-correlation.jsonbefore classifying any synchronized service stop. It contains deterministic, command-free correlation of validated maintenance operation/issue/revision/budget labels. -
Read every
.mdreport under~/dev/claude/review-jobseek-errors/before classifying. The directory name is legacy; keep using it for cross-run continuity unless a migration has already been completed. -
Load prior GitHub issues for dedupe and pattern matching:
gh issue list --label daily-error-review --state all --limit 200 \ --json number,title,state,createdAt,body -
Load recently merged PRs for regression correlation:
gh pr list --state merged --limit 40 --json number,title,mergedAt
Host Signals
Collect host signals before detailed log review. These commands are allowed
without sudo as deploy:
df -h /
df -h /var/lib/docker
free -h
uptime
docker stats --no-stream --format 'table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}\t{{.MemPerc}}'
docker inspect --format '{{.Name}} Image={{.Config.Image}} Created={{.Created}} StartedAt={{.State.StartedAt}} OOMKilled={{.State.OOMKilled}} Status={{.State.Status}} RestartCount={{.RestartCount}} FinishedAt={{.State.FinishedAt}} ExitCode={{.State.ExitCode}} Error={{json .State.Error}}' $(docker ps -aq)
dmesg -T 2>/dev/null | tail -n 500
journalctl -k --since "24 hours ago" --no-pager 2>/dev/null | tail -500
Use journalctl -k as a fallback if dmesg is restricted. If both kernel
log commands are restricted, note the gap under ## Host and do not treat
that gap alone as an incident.
Docker's OOMKilled state is historical and can remain true while a
container keeps running. RestartCount is also scoped to one container
object. Never subtract either value across different pseudonymous container
generations. Treat an OOM or restart as new inside the report window only
when one of these proves it:
- The same pseudonymous
container_generationexists in both evidence snapshots and its cgroup-v2memory.events.oom_killor DockerRestartCountincreased. - A container created inside the window already has a positive counter.
- An exited container finished inside the window with
OOMKilled=true, or kernel logs timestamp the OOM inside the window. - The durable lifecycle stream timestamps the same container's
die,oom,kill, or subsequentstartevent inside the window.
Use lifecycle sequences to classify the mechanism, not to guess a caller:
oomor inspectedoom_killed=truefollowed bydieis a cgroup/kernel OOM.killrecords a Docker API signal; report the signal and say the caller is unknown unless deployment or host evidence identifies it.diewithout a precedingkill/oomis a native process exit; use its exit code and adjacent application logs.stop/die/destroyfollowed by a new container generation'screate/startis a replacement or deployment, whiledie/starton the same generation with an incremented restart count is an automatic restart.
Before treating a Docker API stop or replacement as spontaneous instability, apply the maintenance correlation contract:
- Attribute it only when
host/maintenance-correlation.jsoncontains exactly one validated overlapping operation. Never infer authorization from names, timestamps alone, or an unlabelled one-off. - Report an attributed window under
## Maintenance windowswith operation, tracking issue, reviewed revision, duration, affected services, measured downtime, one-off results, and restoration health. - Do not file/reopen a spontaneous-instability issue for a clean authorized window. Forced termination, OOM/native exit, nonzero one-off, budget overrun, and failed restoration remain actionable; update the maintenance tracking issue or file a deduplicated maintenance-outcome issue.
- Missing, partial, invalid, or conflicting provenance remains unknown. OOM/native exits remain incidents regardless of any adjacent maintenance.
If the lifecycle journal has connect_error/stream_closed records or does
not cover the full window, report the evidence gap explicitly. Do not infer a
clean window from a missing or unavailable journal.
A changed pseudonymous container generation normally means a deployment or
replacement, not a restart-count delta. A sticky OOMKilled=true without a
same-generation counter delta is historical context, not a fresh daily
incident.
Flag an incident if any of these are true:
/usage is at least 85%./var/lib/dockerusage is at least 85% when it is a separate mount.- A generation-safe cgroup or exited-container signal proves an OOM inside the window.
- A same-container
RestartCountincrement proves a restart inside the window. - An unattributed lifecycle
die/startsequence proves a restart inside the window even when yesterday's snapshot is unavailable. An attributed sequence follows the maintenance-outcome rules above. - 15-minute load average is greater than CPU count x 2.
- Swap usage is greater than 50% of swap total.
- Kernel logs show
Out of memory: Killed process,I/O error,EXT4-fs error,Call Trace:, or systemd OOM-killer entries. - Any expected long-running service is missing from
docker ps.
Log Collection
For each long-running service:
docker logs --since "<ISO>" --until "<ISO>" <container> 2>&1
Parse structlog JSON where possible. Fall back to line matching only for
non-JSON lines. Extract level, event, service_name, exception class,
and a stable stack or message fingerprint.
For exited ephemeral containers inside the window:
docker logs <id> 2>&1 | tail -n 1000
Attribute those errors to a synthetic service derived from the command, such
as refresh-typesense or notify-indexnow.
Classification
Group errors by (service, exception class, stable message stem).
known: appears in any prior daily report within 14 days. Count it but do not file.novel: absent from every prior report.regression: a known class that was absent for at least 3 consecutive days and is back today.spike: a known class whose 24-hour count is at least 3x its 7-day median. Require at least 3 prior days with non-zero signal.incident: host-signal trigger, unexpected zero logs from a long-running service for more than 1 hour, data loss or corruption signal, repeated OOM kills, or exporter/drain stall with queue depth climbing and no progress.
Incomplete evidence narrows confidence but does not make GitHub filing
impossible. If a bundle covers less than the full 24-hour window, classify
full-window trends as unclassified — partial window, but still file or
update an issue when the available evidence shows a concrete, redacted,
deduped error class that would independently meet a filing criterion inside
the observed window. In that case, state the observed window and evidence gap
in the report and issue body, avoid unsupported 24-hour trend claims, and use
severity based only on the verified blast radius.
Report
Always write a report, even on a healthy day.
Path:
~/dev/claude/review-jobseek-errors/YYYY-MM-DD.md
If today's file already exists, append a ## Rerun HH:MM UTC section instead
of overwriting it.
Report schema:
# Daily error review - YYYY-MM-DD
Window: YYYY-MM-DD HH:MM UTC -> YYYY-MM-DD HH:MM UTC
## Host
| metric | value |
|---|---|
## Totals
| service | info | warning | error |
|---|---:|---:|---:|
## Error classes (24h)
| class | service | count | status |
|---|---|---:|---|
## Maintenance windows
| operation / issue | revision | duration | affected services | outcome |
|---|---|---:|---|---|
## Details
## Filed issues
## Health
Under ## Details, include every novel, regression, spike, or incident class:
- Error class and service.
- 24-hour count plus first-seen and last-seen UTC.
- One redacted sample log line.
- Brief root-cause hypothesis with 2 or 3 candidates. Use file paths only.
- Recent PR cross-references when correlation is plausible.
Under ## Maintenance windows, include every deterministic correlation row,
including clean authorized operations. Report forced termination, OOM/native
exit, failed/nonzero one-offs, budget overruns, measured service downtime, and
failed restoration explicitly. Use none when no maintenance window is
present.
Under ## Filed issues, list issue URLs filed or updated in this run, or
write none.
GitHub Issues
File or update issues for novel, regression, spike, or incident
classes. Partial-window evidence is allowed when the observed error class is
concrete and issue-worthy on its own; caveat the reduced window in the signal
section rather than suppressing the issue categorically.
Deduplicate first:
gh issue list --label daily-error-review --state all \
--search "<error class stem or keyword>"
Match by error class and service, not exact title.
- If a matching open issue exists, add a comment with today's window, count, first-seen, and last-seen. Do not open a duplicate.
- If a matching closed issue from the last 30 days exists, reopen it with a comment explaining the recurrence.
- Otherwise open a new issue.
New issue template:
## Summary
One paragraph with service, what broke, whether it is visible or silent, the
explicit 24-hour window, and suspected PR reference if any.
## Signal
5-day trend table (-4, -3, -2, -1, today) of count or batched log-line count.
Explain the unit.
## Sample log line
```json
{"level":"error","event":"redacted sample"}
```
## Root cause hypothesis
1. Candidate with file-path reference only and fix direction.
2. Candidate with file-path reference only and fix direction.
## Impact
User-facing consequence such as stale data, disabled boards, or silent
failure. If invisible beyond `level=error`, say so.
Title shape:
[daily-error-review] <service>.<error_stem> (<trend>; <blast>)
For critical incidents, prefix the title with [critical].
Labels are required:
daily-error-review- Exactly one of
error-review:critical,error-review:high,error-review:medium, orerror-review:low
Severity mapping:
error-review:critical: active incident or data-loss risk.error-review:high: regression or full-service blast radius.error-review:medium: spike or partial-service regression.error-review:low: novel but contained.
Redaction
The repository is public. Redact before any GitHub write:
- Full URLs with query strings: keep host and path shape only.
- IPs, port numbers, and hostnames beyond
jseek.coor public CDNs. - UUIDs, emails, user IDs, and company internal IDs.
- Raw Docker/container IDs and cloud resource IDs. Use the bundle's
pseudonymous
container_generationonly in the private report when needed. - JWTs, cookies, API keys, and Authorization headers.
- SQL literal values; keep statement skeleton only.
- Request or response bodies beyond the exception message itself.
- Absolute paths containing
/homeor/root. - Stack traces with parameter values. Keep exception class and innermost frame
as
filename:functiononly.
Local reports under ~/dev/claude/review-jobseek-errors/ are private. Raw
details belong there, not in GitHub issues.
Guardrails
Allowed on the remote box:
docker logsdocker psdocker inspectdocker stats --no-streamdffreeuptimedmesgjournalctl -kwhen accessible- Reading files under
/home/deploy/ - Reading and writing temporary files under
/tmp/that you created this run
Forbidden on the remote box:
- Restarts, stops,
docker exec,docker kill, and anydocker composesubcommand. systemctl.- File edits outside
/tmptemporary files you created this run. - Deletions outside
/tmptemporary files you created this run. - Any
sudo. - Any outbound HTTP to third-party hosts.
Writing scripts to /tmp on the remote box is allowed only when the script is
idempotent and read-only. Remove scripts after use.
If a command's effect is not completely known to be read-only, stop, write the intent to today's report, classify the gap as an incident, and do not run it.
Escalation
For any incident:
-
File or update an issue labelled
daily-error-reviewanderror-review:critical. -
Put
[critical]at the start of a new critical issue title. -
Append a line to
~/dev/claude/review-jobseek-errors/ALERTS.md:YYYY-MM-DD HH:MM UTC <one-sentence> <issue URL> -
Do not attempt to fix production during this routine.
Verification Rollout
Use this rollout when changing the routine or moving it to a new automation surface:
- Read-only dry run: collect host signals, logs, prior reports, and GitHub issue state; write the report; list would-file or would-update issues without creating, reopening, or commenting on GitHub issues.
- Production pilot: run against the real host and real GitHub issue state, but file only when the criteria are unambiguous and the evidence is already redacted.
- Require two clean runs before normal filing behavior. A clean run means the report is written, known issues are deduped, no forbidden command is used, and any GitHub write is justified by the classification rules.
- After edits, scan docs and agent instructions for stale primary-runtime wording such as Claude-only phrasing, direct API billing, or obsolete command names.
Version History
- 7517cfd Current 2026-08-02 21:45


