performance
GitHub用于静态Web服务器性能分析与优化,涵盖延迟、吞吐量及资源使用。指导通过profiling定位瓶颈,提供Rust代码优化(如避免分配、预计算)及HTTP连接处理策略,确保满足性能目标。
Trigger Scenarios
Install
npx skills add static-web-server/static-web-server --skill performance -g -y
SKILL.md
Frontmatter
{
"name": "performance",
"description": "Optimize or review performance for the Static Web Server (SWS) project — profiling, bottlenecks, resource usage, compression, and caching"
}
Performance Optimization
Load this skill when profiling, optimizing, or reviewing code for performance — latency, throughput, memory, or CPU.
When to load: a change touches the request hot path (handler.rs, static_files/, compression.rs, fs/stream.rs), a benchmark is being added or interpreted, a regression in requests/sec or latency is suspected, or a new dependency may affect binary size or runtime cost.
General Approach
- Measure before optimizing: Use profiling tools (perf, flamegraph) to identify bottlenecks. Never optimize based on intuition
- Set a target: Define acceptable latency/throughput before starting. Stop optimizing when the target is met
- Optimize the hot path: Focus on code that runs on every request. Startup code and config parsing are low priority
Rust Performance
- Profile with
perfandflamegraph:cargo flamegraph --bin static-web-serverfor CPU profiles. Profile under load (e.g.,wrkorbombardier) - Profile heap allocations with DHAT: Use DHAT or dhat-rs to find hot allocation sites. Reducing 10 allocations per million instructions can have measurable impact
- Avoid unnecessary allocations: Prefer
&Pathover&PathBuf,&[u8]overVec<u8>, pass by reference where ownership is not needed - Pre-compute at startup: Canonicalize paths, parse config, compile regex patterns, build Aho-Corasick automata once. Never on the request path
- Pre-allocate collections: Use
Vec::with_capacity,String::with_capacitywhen the size is known - Stream large responses: Use
tokio::fs::File+tokio::io::copyfor file serving. Never buffer the full file in memory - Inline small hot functions: Use
#[inline]on small functions called on every request (e.g., header name normalization, MIME lookups). Use#[cold]on error-path functions to guide branch prediction away from the hot path - Prefer
filter_mapoverfilter().map(): Avoids an intermediate layer in hot iterator chains - Use
iter().copied()for small types: When iterating over&u8,&u32, etc.,.copied()lets LLVM generate better code than receiving references - Use
chunks_exactwhen chunk size evenly divides length: Faster thanchunksbecause it eliminates a remainder check per iteration - Prefer
ok_or_elseoverok_or:ok_or(expensive())always evaluates its argument.ok_or_else(|| expensive())is lazy and only evaluates onNone - Eliminate bounds checks in hot loops: Use iteration instead of index-based access, or add an upfront assertion on the range to let the compiler prove bounds are safe
HTTP Performance
Connection Handling
- HTTP/1.1 keep-alive: Enabled by default via Hyper. Reduces connection setup overhead for subsequent requests
- HTTP/2 multiplexing: Enable with
--http2 --tls. Multiple concurrent streams over a single TCP connection - Worker threads: Default is
num_cpus * 1. Increase--threads-multiplierfor workloads with mixed CPU and I/O blocking (e.g., many concurrent clients with dynamic compression enabled — compression per-request is CPU-bound but high concurrency adds I/O wait interleaving). For pure CPU-bound workloads with minimal I/O, increasing threads beyond CPU count rarely helps. - Max blocking threads: Default 512. For I/O-heavy patterns (large file serving), this is sufficient
- Graceful shutdown: Use
--grace-periodto allow in-flight requests to complete before shutdown
Compression Tradeoffs
- Static compression is free: Pre-compressed
.br/.gz/.zstfiles are served with zero CPU. Always prefer this for production - Dynamic compression overhead: On-the-fly compression trades CPU for bandwidth. Use
--compression-level fastestfor high-traffic sites - Minimum size threshold: Responses below 860 bytes skip dynamic compression entirely — the overhead exceeds any bandwidth savings
- Compression algorithm priority (by compression ratio × speed): zstd > brotli > gzip > deflate. zstd offers the best ratio-speed tradeoff
Caching Headers
- Cache-Control is enabled by default: SWS sets
max-agebased on file extension:- 1 year for static assets (
.css,.js,.png,.woff2, etc.) - 1 hour for feeds/API (
.json,.xml,.rss,.atom) - 1 hour fallback for unknown extensions
- 1 year for static assets (
You can override these defaults per file or extension using the configuration file. The above values are defaults, not hardcoded limits.
- Conditional requests: SWS supports
If-Modified-SinceandIf-Unmodified-SinceviaConditionalHeaders. Returns 304 when the file hasn't changed - ETag not implemented: SWS uses
Last-Modifiedinstead. For byte-level cache validation, put SWS behind a CDN or reverse proxy
File I/O Performance
Buffering
- Optimal buffer size:
optimal_buf_size()selects the best buffer size based on file metadata (usesstd::fs::Metadata::blksize()when available) BufReaderwithtake(): For byte-range requests, aBufReaderwraps the file handle and limits bytes read to the requested range- Streaming avoids full-file buffering:
FileStreamreads in chunks. The response body is a stream, not a byte buffer
Path Operations
- Canonicalize once at startup: The root directory is canonicalized in
server/opts.rs. Per-request path resolution reuses this try_metadata()caches nothing: Each call (insrc/fs/meta.rs) is a filesystem syscall. The experimental memory cache feature (mini-moka, insrc/mem_cache/) caches file metadata and content- Avoid
clone()in the hot path:static_files.rsavoids cloning file paths for non-directory requests
Pre-compressed Static Files
- Zero-CPU serving: SWS detects
.br/.gz/.zstvariants viaAccept-Encodingand serves them directly. No compression step runs - Build-time pre-compression: Generate variants with maximum quality:
brotli -q 11,gzip -9,zstd -19. SWS serves them as-is - Vary header:
Vary: Accept-Encodingis appended so caches know to store multiple variants
Memory
- Minimal per-connection state: SWS stores only the remote address and handler opts (shared via
Arc). No per-connection buffers - Response body is a stream: File contents are streamed, not buffered. Exception: small generated responses (health endpoint, error pages, directory listing HTML)
- Experimental in-memory cache:
mini-moka(insrc/mem_cache/) caches hot files in memory with LFU admission and LRU eviction. Configurablecapacity(default 100 entries),ttl(default 1800s),tti, andmax_file_size. Keys useCompactStringto reduce allocation
Allocation Patterns
- Prefer
clone_fromover reassign-and-clone:a.clone_from(&b)reusesa's existing heap allocation when possible, avoiding an extra alloc/free. Especially valuable forVecandStringin hot loops - Reuse collections across iterations: Declare the collection outside the loop, call
.clear()at the end of each iteration. Avoids repeated alloc/free while keeping the heap allocation alive - Use
Cow<'_, str>/Cow<'_, Path>for mixed borrowed/owned data: Avoids allocating aString/PathBufwhen the data is already a static literal or an existing slice that won't be modified - Use
SmallVec<[T; N]>for short, stack-like sequences: When most allocations hold ≤ N elements (e.g., header value lists, index file candidates),smallvecavoids heap allocation entirely for the common case - Convert finalized
VectoBox<[T]>withinto_boxed_slice(): Drops the unused capacity word, shrinking the type from 3 words to 2. Good for config-time data that is built once and never grown - Return
impl Iterator<Item=T>instead ofVec<T>from helpers: Avoids an allocation when the caller only needs to iterate - Avoid
format!when a literal orwrite!suffices: Everyformat!call allocates aString. Write directly to a&mut Stringor usestd::fmt::Writeinstead
Type Sizes
- Keep hot types under 128 bytes: The compiler emits
memcpyfor values larger than 128 bytes. If a hot type exceeds this, check its layout withRUSTFLAGS=-Zprint-type-sizes cargo +nightly build --release - Box large enum variants: If one variant is much larger than the others, box its fields to bring all variants to a similar small size. Reduces stack pressure and cache churn
- Use smaller integer types for index/count fields: Prefer
u32overusizefor counts and offsets stored in frequently instantiated structs (e.g., header tables, path segments). Cast tousizeat use sites - Assert type sizes in tests: Add
static_assertions::assert_eq_size!(HotType, [u8; N]);for performance-critical types so that accidental size regressions cause a compile error
Release Build Configuration
The default cargo build --release profile is a good starting point, but the following options can improve throughput for production SWS builds:
| Option | Effect | Cargo.toml |
|---|---|---|
lto = "thin" |
Cross-crate inlining, 5–15% speedup, moderate compile cost | [profile.release] |
codegen-units = 1 |
Single codegen unit, enables more optimizations, slower compile | [profile.release] |
panic = "abort" |
Removes unwinding machinery, smaller binary, slight speedup | [profile.release] |
For a custom server build where broad CPU compatibility is not required:
RUSTFLAGS="-C target-cpu=native" cargo build --release
This emits AVX/SSE instructions optimal for the build machine, which can improve compression throughput.
Note:
target-cpu=nativeproduces a non-portable binary. Do not use for distributed release artifacts.
Benchmarking
Micro-benchmarks (CodSpeed)
The benches/ directory is a standalone crate with Criterion benchmarks for hot-path
functions (path sanitization, header handling, redirects, basic auth). They run on every
pull request via the codspeed workflow and are tracked on
CodSpeed.
cd benches
# Plain Criterion run (local timings only)
cargo bench
# Same suite measured with CodSpeed CPU simulation (deterministic, flamegraphs)
cargo codspeed build
codspeed run --mode simulation -- cargo codspeed run
Add a bench when a change touches the request hot path so a future regression is caught by CI instead of by users.
Tools
- HTTP load generators:
wrk,bombardier,oha,hey - CPU profiling:
perf record+flamegraph,cargo flamegraph,cargo instruments(macOS),samply(cross-platform) - Allocation profiling:
dhat-rs(all platforms), DHAT via Valgrind (Linux) — identifies hot allocation sites - Monitoring: Prometheus metrics via
--metrics+/metricsendpoint
What to Measure
- Requests per second at concurrency levels: 1, 10, 100, 1000
- Latency percentiles: p50, p95, p99
- Memory usage: RSS before and under load
- CPU utilization: Per-core usage during load test
- Allocation rate: Use DHAT to confirm per-request allocations are not growing unexpectedly
Checklist
- Is there a benchmark or profile showing the bottleneck?
- Are paths canonicalized once, not per-request?
- Is static compression used where possible (zero CPU)?
- Is dynamic compression size-threshold applied (860 bytes)?
- Are large files streamed, not buffered?
- Are Cache-Control headers set appropriately for the content type?
- Are worker threads configured for the workload?
- Is keep-alive or HTTP/2 enabled for connection reuse?
- Are hot allocation sites identified with DHAT or similar?
- Are collections pre-allocated or reused rather than recreated per request?
- Are
clone()calls on the hot path justified — or replaceable withclone_from,Cow, or a reference? - Are hot types under 128 bytes (no unintended
memcpy)? - Are
#[inline]/#[cold]attributes applied where profiling shows they help? - Are
ok_or_else/ lazy combinators used instead of eager alternatives in hot paths?
Version History
- 21dc11b Current 2026-08-20 17:42


