Fail closed
Unknown names should not fall through to a real application. Default servers exist to reject ambiguity, not to become accidental catch-alls.
A practical field guide for internet-facing nginx: split protocol risk before application routing, reject what should never arrive, bound what can become expensive, keep trust explicit, use Lua as a small policy layer, and make the remaining behavior observable. No WAF theatre. No giant signature dump. No claim that one config fits every application.
The useful hardening decisions are usually boring: know which hosts, methods, protocols and request sizes you intend to serve, then make everything else explicit.
Unknown names should not fall through to a real application. Default servers exist to reject ambiguity, not to become accidental catch-alls.
Requests consume sockets, buffers, upstream time and application work. Put sensible ceilings around those resources before you need them.
A static site does not need POST, CONNECT or TRACE. An upload endpoint does not need the same body limit as the rest of the site.
Forwarded addresses, scheme headers and proxy chains are security boundaries. Treat them as configuration, not ambient truth.
Request IDs, status, latency and upstream timing usually tell you more than indiscriminate request dumps — and are much safer to retain.
Dry-run limits where possible, inspect real traffic, then enforce. Hardening that breaks legitimate clients is just another outage.
The strongest controls are usually not giant blocklists. They make expensive work conditional, separate protocol risk from application policy, keep trust boundaries explicit and reject impossible traffic before it reaches application code.
If a request is unambiguously impossible for the service, dispose of it before canonical redirects, proxying, sessions or application bootstrap consume work.
A cheap static 404 and an expensive authenticated render do not need the same budget. Put the tighter control around the operation that actually burns CPU, DB time, sockets or third-party calls.
IP addresses, user agents, TLS fingerprints and referrers are useful signals, not identity. When possible, bind enforcement to the costly resource, route or application state instead.
nginx rate-limit zones can use an empty key to skip unrelated requests. That makes a narrow login, callback or expensive-route budget possible without taxing the whole site.
Telemetry may combine many weak signals to explain behavior. Enforcement should still use a small, reviewable contract and should not quietly promote a heuristic into an authentication decision.
If a webhook, upload client or legacy protocol genuinely needs an exception, scope it to the exact host, route, method and budget instead of weakening the global edge.
UDP/443 has its own listener and failure modes. Give unknown QUIC traffic a fail-closed default too; do not assume the TLS/TCP default server automatically covers it.
Authentication helpers and subrequest services do not need public reachability. Keep them on loopback or a private socket and expose only the application boundary that consumes their result.
Keep a compact default format for correlation, optionally split error-class requests into a second stream, and decide separately when deeper debug capture is justified.
One production pattern that has held up well is to split TCP/443 by ALPN before TLS termination: HTTP/2 enters a clean lane, HTTP/1.x enters a constrained compatibility cage, and HTTP/3 terminates separately on UDP/443 only where it is intentionally enabled.
The front listener only prereads enough TLS metadata to see ALPN. It does not terminate TLS here.
ssl_preread selects a loopback lane and PROXY protocol preserves the client address for the terminating server.
TLS terminates here. Shared application policy still applies, but the transport does not inherit HTTP/1 compatibility penalties.
Legacy clients, control-plane tools and deliberate compatibility traffic still need a defined policy.
Traffic without the clean-lane ALPN is routed to a separate TLS listener instead of silently receiving the same policy.
Keepalive can be disabled, timeouts shortened, connection counts bounded and expensive routes given tighter budgets.
HTTP/3 is a different transport edge. The TCP shunt cannot protect or route it.
Only selected public application hosts advertise and terminate h3. Unknown QUIC traffic deserves its own fail-closed default.
Application authorization remains the same, while transport-specific limits, logging and rollout stay independently testable.
The stream layer can inspect ALPN from ClientHello without terminating TLS. That lets the HTTP servers remain normal TLS endpoints while still receiving distinct protocol classes and the original client address.
stream {
map $ssl_preread_alpn_protocols $lane {
~\bh2\b 127.0.0.1:8443;
default 127.0.0.1:9443;
}
server {
listen 443 reuseport;
ssl_preread on;
proxy_protocol on;
proxy_pass $lane;
}
}
A constrained H1 lane can keep old clients working without giving them the same socket lifetime and expensive-route budget as the preferred transport. The point is not “HTTP/1 is evil”; the point is that compatibility can have a smaller blast radius.
keepalive_timeout 0;
client_header_timeout 5s;
client_body_timeout 5s;
send_timeout 5s;
reset_timedout_connection on;
limit_conn h1_clients small_limit;
# Give expensive H1 routes their own tighter budget.
location ~ expensive_dynamic_route {
limit_req zone=h1_expensive burst=small_burst nodelay;
try_files $uri $uri/ /index.php?$args;
}
A useful file split is common-start → lane-specific policy → common-end. The common start restores the trusted client address and applies global request invariants; the lane file contains protocol-specific controls; the common end defines exact application endpoints, Lua access policy, static handling and logging.
server {
listen 8443 ssl proxy_protocol;
http2 on;
include edge/common_start.conf;
include edge/h2_lane.conf;
include edge/common_end.conf;
}
server {
listen 9443 ssl proxy_protocol;
include edge/common_start.conf;
include edge/h1_cage.conf;
include edge/common_end.conf;
}
Browsers and modern crawlers may prefer h2 or h3, while inspection tools, provider control planes and legacy clients can still arrive over H1. Keep the compatibility exception conjunctive: protocol + verified network/client class + expected host + allowed method/route.
map "$is_h1:$verified_crawler:$service_h1_ok" $allow_h1 {
default 0;
~^1:1: 1;
~^1:0:1$ 1;
}
# The exception opens the protocol lane only.
# Method, body, host, rate and app authorization still apply.
Real edge policy usually needs more than one kind of exception. The useful pattern is to keep those classes separate: verified crawlers, fixed operational networks, route-bound service clients and weak telemetry signals should not collapse into one magic whitelist bit.
A crawler can be treated as a low-risk exception signal because it arrives from a maintained provider network, because it is a narrowly recognized service client, or both. That can be enough to exempt it from a weak heuristic — but it is still not authentication.
When the exception itself is security-relevant — for example allowing an older protocol lane — combine the crawler claim with the matching provider CIDR. The same source can therefore be “good enough to skip a heuristic” yet still fail a stricter compatibility gate.
Webhooks, inventory fetchers and preview clients often have recognizable request shapes. Treat that as an endpoint contract: exact route, method, content type, budget and application-level verification — not as a blanket trusted-agent rule.
ASN, country, referrer, Accept-Language, browser hints and TLS characteristics are useful for correlation. They may explain a request or contribute to a narrow gate, but they should not silently become authentication.
nginx geo maps are a convenient way to consume maintained CIDR files. In production I use separate generated files for public crawler networks such as Apple, Bing, DuckDuckGo, Google, OpenAI and Qwant, then validate the whole config before reload.
geo $crawler_net_ok {
default 0;
include /etc/nginx/cidrs/applebot.cidr;
include /etc/nginx/cidrs/bingbot.cidr;
include /etc/nginx/cidrs/duckduckbot.cidr;
include /etc/nginx/cidrs/googlebot.cidr;
include /etc/nginx/cidrs/google-special.cidr;
include /etc/nginx/cidrs/openai-search.cidr;
include /etc/nginx/cidrs/openai-user.cidr;
include /etc/nginx/cidrs/qwantbot.cidr;
}
A public CIDR says where the request came from. The user agent says which crawler role it claims. For stronger protocol or compatibility exceptions, require both. For lower-stakes heuristic exemptions, a separate broader goodbot signal can remain intentionally less strict.
map $http_user_agent $crawler_ua_claim {
default 0;
~*Googlebot 1;
~*bingbot 1;
~*DuckDuckBot 1;
~*Applebot 1;
~*OAI-SearchBot 1;
~*ChatGPT-User 1;
}
map "$crawler_ua_claim:$crawler_net_ok" $verified_crawler {
default 0;
"1:1" 1;
}
# A broader low-risk signal can deliberately be OR-based.
map "$crawler_net_ok:$service_ua_ok" $goodbot_signal {
default 0;
~^1: 1;
~:1$ 1;
}
Do not make one global $whitelist mean “skip security.” Use separate variables and combine them only where the endpoint actually needs the exception. Documentation-range addresses are used below deliberately.
geo $ops_net {
default 0;
192.0.2.10/32 1;
198.51.100.0/28 1;
}
map "$verified_crawler:$ops_net" $crawl_policy_exempt {
default 0;
~^1: 1;
~:1$ 1;
}
# A payment callback would use its own route-bound rule instead.
Generated crawler lists should come from official feeds, be rendered into deterministic nginx include files, and be replaced per-file atomically. The updater should stop on fetch or parse errors and reload only after nginx -t succeeds. Staging the entire set as one transaction is an additional improvement when you need all feeds to advance together.
#!/bin/sh
set -eu
tmpdir=$(mktemp -d)
trap 'rm -rf "$tmpdir"' EXIT
# fetch + validate provider JSON
# render: 203.0.113.0/24 1;
install -m 0644 "$tmpdir/provider.cidr" /etc/nginx/cidrs/provider.cidr.new
mv /etc/nginx/cidrs/provider.cidr.new /etc/nginx/cidrs/provider.cidr
nginx -t
service nginx reload
The fast path should stay mostly native nginx: map, geo, exact locations and native limit zones. Lua earns its place when a small amount of stateful or structured policy is clearer there than in another fifty maps.
Cheap map/geo variables derive protocol, route, crawler, whitelist and cost classes.
H1/H2/H3 transport policy removes impossible traffic before application routing.
One access-phase dispatcher executes an explicit ordered module chain once per request.
Native request/connection zones count only the routes or resources that are actually expensive.
Authentication, payment truth and business authorization remain application contracts.
Log-phase policy records the class and outcome without changing an already completed request.
Guard the dispatcher in ngx.ctx. A rewrite or internal redirect should not accidentally execute the access chain twice.
Use an allowlist of known module names. Configuration should select from reviewed policy modules, not dynamically require arbitrary Lua paths.
Each policy module gets one run(ctx)-style contract and a tiny set of valid return types. Ambiguous results are errors, not magic.
Missing modules, bad chains, duplicate modules and runtime exceptions should become a controlled edge failure — not a silent bypass.
Do not parse every body and flatten every header on every request. Gate expensive diagnostic capture to the small class of requests that actually need it.
Use the log phase for correlation and enrichment. Security decisions must already be complete before response logging begins.
A compact dispatcher can normalize an ordered chain, reject duplicates, require an explicit module allowlist, verify the module contract and fail closed on load/runtime errors. That is much easier to reason about than scattered access_by_lua_block fragments.
-- access.lua (simplified)
if ngx.ctx.edge_dispatched then return end
ngx.ctx.edge_dispatched = true
local allowed = {
collect_context = true,
route_policy = true,
}
for name in configured_chain():gmatch("([^,]+)") do
if not allowed[name] then
return fail_closed("module_not_allowed", name)
end
local ok, mod = pcall(require, name)
if not ok or type(mod.run) ~= "function" then
return fail_closed("module_contract_invalid", name)
end
local ran, result = pcall(mod.run, ngx.ctx)
if not ran then
return fail_closed("module_runtime_failed", name)
end
if result == true then return end
end
Large map/geo sets are lazy-evaluated by nginx. They are excellent for composing protocol, path, CIDR and method signals into a small number of policy variables. Lua should receive the already-classified request rather than reimplementing the whole edge parser.
geo $provider_net {
default 0;
include /etc/nginx/cidrs/provider.cidr;
}
map $request_uri $is_expensive_route {
default 0;
~^/search 1;
}
map "$is_expensive_route:$provider_net" $access_chain {
default "collect_context,route_policy";
"0:1" "route_policy";
}
Detailed header/body capture is operationally useful, but expensive and sensitive. Derive a small $is_deep class from route/policy signals, capture only there, truncate bodies, and pass the minimum context through ngx.ctx to the log phase.
map "$is_asset:$verified_crawler:$interesting_request" $is_deep {
default 0;
"0:0:1" 1;
}
# access phase: normalize + truncate only when $is_deep = 1
# log phase: serialize the prepared context, no new security decision
Separate streams for normal traffic, verified crawlers, rate limits, high-confidence sinks and deep diagnostics make investigations faster. A higher-confidence event can suppress duplicate lower-value log lines so one request does not masquerade as five incidents.
map $status $is_429 { default 0; 429 1; }
map $policy_sink $log_policy { default 0; ~.+ 1; }
access_log /var/log/nginx/access.json edge_json if=$log_normal;
access_log /var/log/nginx/policy.json edge_json if=$log_policy;
access_log /var/log/nginx/rate.json edge_json if=$is_429;
nginx is a phase machine. A readable edge design follows those phases deliberately: transport admission first, address normalization before policy inputs, native cost controls before application access, application truth after edge heuristics, and logging only after the outcome is known.
TCP/UDP, ALPN preread and protocol-lane selection happen before the HTTP request exists.
Restore the trusted client address before downstream CIDR, rate and audit policy consumes it.
Server/location rewrite logic can reject or transform early, before access and content work.
Native connection and request limiters execute here. The expensive request may die before access policy.
Auth modules and access-phase Lua make narrow allow/deny decisions from already-normalized inputs.
try_files and related routing decide whether static content, an application entry point or a miss is next.
Static output, FastCGI, proxying or a content handler executes. Payment/auth/business truth belongs here or deeper.
Enrich and route telemetry after the response decision. Log phase should never retroactively become enforcement.
Two directives next to each other in a file may execute in different nginx phases. Review the module hook, not just the line number.
Real client address, canonical path and protocol class should be stable before they become CIDR, limiter or logging inputs.
Use native preaccess limits for simple cost ceilings. Do not wake a richer policy engine merely to rediscover that a client exceeded a counter.
An H1 compatibility exception grants H1. A crawler exception skips a weak crawler heuristic. Neither implies application authorization.
UA, CIDR, fingerprint and referrer can alter edge policy. They must not silently upgrade payment, account or session truth.
Once an impossible host, method, protocol or route is rejected, a later generic fallback should not resurrect it.
A common split such as common-start → lane policy → common-end improves ownership, but the runtime still follows nginx phases. In a real edge, real-IP normalization can execute before a limiter even if its directive lives in another include; limit_req/limit_conn run in preaccess; access Lua runs later; try_files later again; log handlers run last.
# transport plane
accept -> preread -> lane
# HTTP plane
post-read -> trusted client address
rewrite -> early normalization / rejection
preaccess -> native req/conn budgets
access -> auth + Lua policy
precontent -> try_files / route resolution
content -> static / FastCGI / proxy
log -> classify + correlate outcome
# config-file order is not a substitute for this model
Keep lane admission, heuristic exemption, resource budgeting and application authorization as different outputs. This makes policy reviewable: one signal can relax one heuristic without mutating unrelated controls.
map "$protocol:$verified_service" $lane_allowed {
default 0;
"h2:0" 1;
"h3:0" 1;
"h1:1" 1;
}
map "$verified_crawler:$route_class" $heuristic_exempt {
default 0;
"1:public" 1;
}
# application authorization is intentionally not derived here
A matrix makes exceptions auditable. The examples below are deliberately generic, but mirror a useful production principle: crawler, operator, callback and browser classes may receive different transport or heuristic treatment while method, body, resource and application contracts stay independently enforced.
| Traffic class | Protocol admission | Edge exemptions | Cost budget | Application truth | Telemetry |
|---|---|---|---|---|---|
| Normal browser | h2/h3 normalh1 cage | None by default. | Route/resource-specific native limits. | Normal session, auth, CSRF and business rules. | Normal access stream; deep capture only when selected. |
| Verified crawler | h2/h3h1 only if compatibility policy allows | May skip crawler-sensitive heuristics after provider verification. | Still finite; crawler-specific budgets may differ. | No account/payment authorization upgrade. | crawler stream for volume and drift. |
| Operator / service allowlist | Only the protocols the service actually needs. | Narrow, named exemptions — never a universal bypass. | Dedicated ceiling appropriate to the service. | Service authentication and route authorization still required. | allowlist stream with explicit reason class. |
| Provider callback | Endpoint-bound compatibility, including legacy H1 only when required. | Browser heuristics may be irrelevant on that exact endpoint. | Dedicated method/body/rate envelope. | provider signature/assert remains authoritative. | Callback outcome + retry correlation; no secret material. |
| Internal auth subrequest | not public internal location / loopback. | Public crawler/whitelist policy does not apply. | Small bounded internal budget. | Auth service returns the authorization decision. | Minimal auth result/latency; no credentials. |
| Unknown / weak H1 client | cage or reject | None. | Tight connection, timeout and route budgets. | Same application checks if it reaches content. | Protocol/rejection class for compatibility analysis. |
| Repeated expensive resource | Any otherwise-valid lane. | No identity shortcut from IP/UA churn. | state + canonical resource where appropriate. | Business state remains unchanged by rate classification. | 429 / cost stream with bounded context. |
These snippets are examples, not a drop-in complete config. Values are intentionally readable starting points; application requirements win.
Hide the nginx version, bound headers and bodies, and set finite read/write timeouts. Large uploads or long-running upstreams should get narrowly scoped exceptions instead of globally relaxed defaults.
server_tokens off;
client_max_body_size 1m;
client_body_buffer_size 128k;
client_header_buffer_size 2k;
large_client_header_buffers 4 16k;
client_header_timeout 10s;
client_body_timeout 15s;
send_timeout 30s;
keepalive_timeout 30s;
TLS 1.2 + 1.3 remains a sensible compatibility baseline. Keeping 0-RTT disabled is conservative unless the application and every relevant request path are designed around replay semantics.
ssl_protocols TLSv1.2 TLSv1.3;
ssl_session_cache shared:SSL:20m;
ssl_session_timeout 1d;
ssl_session_tickets off;
ssl_early_data off;
A dedicated TLS default server can reject handshakes for names you do not host. This is cleaner than letting an unknown SNI select whichever application server happens to be first.
server {
listen 443 ssl default_server;
server_name _;
ssl_reject_handshake on;
}
For a truly static vhost, other methods are noise. The return-only conditional below is intentionally simple; dynamic applications should define their method policy per endpoint instead.
map $request_method $static_method_denied {
default 1;
GET 0;
HEAD 0;
}
server {
# A return-only if is easy to reason about.
if ($static_method_denied) {
return 405;
}
location / {
try_files $uri $uri/ =404;
}
}
Reject hidden paths unless you deliberately publish one. If you use ACME HTTP-01 or a well-known metadata path, carve that path out explicitly rather than weakening the general rule.
location ~ /\. {
return 404;
}
nginx has native request and connection budgets. Start with broad, measurable ceilings and narrow them only where the workload justifies it. If nginx sits behind another proxy, establish the trusted real-client address first or an IP-keyed limit is meaningless.
For HTTP/2 and HTTP/3, remember that nginx counts each concurrent request as a connection for limit_conn.
limit_conn_zone $binary_remote_addr zone=per_ip_conn:10m;
limit_req_zone $binary_remote_addr zone=public_req:10m rate=20r/s;
limit_conn_status 429;
limit_req_status 429;
server {
limit_conn per_ip_conn 32;
limit_req zone=public_req burst=60 nodelay;
}
A login or sign-up budget should normally count only those requests. Build the key from the request shape you care about and return an empty value for everything else. nginx does not account empty keys in the zone.
map "$request_method:$uri" $login_limit_key {
default "";
~^POST:/user/(?:login|sign-up)$ $binary_remote_addr;
}
limit_req_zone $login_limit_key zone=login:1m rate=10r/m;
server {
limit_req zone=login burst=10 nodelay;
}
Per-IP controls are useful coarse ceilings, but they are often the wrong identity model. For stateful applications, a tighter budget can follow a server-issued session plus a canonical resource path so rotating egress addresses do not reset the cost accounting.
The exact key is application policy. Keep it out of generic public examples if it would reveal bypass-sensitive internals.
# Coarse edge ceiling:
client address -> broad connection/request budget
# Narrow expensive-operation ceiling:
server-issued state + canonical resource
-> cost budget for that exact operation
# Do not use as identity by itself:
UA / ASN / referrer / TLS fingerprint
Subrequest-based authentication works well when the helper is reachable only from nginx. Keep the auth location internal and the service on loopback or a private Unix socket; the public application consumes only the allow/deny result.
location = /_auth {
internal;
proxy_pass http://unix:/run/app-auth.sock;
proxy_pass_request_body off;
proxy_set_header Content-Length "";
proxy_set_header X-Original-URI $request_uri;
}
location /private/ {
auth_request /_auth;
proxy_pass http://app_backend;
}
Payment providers, webhooks, monitoring agents and old package clients sometimes need a protocol, method or timeout that normal browser traffic does not. Keep the exception on the exact route and preserve the stricter default everywhere else.
location = /provider/callback {
# Only the provider contract gets the exception.
limit_req zone=callbacks burst=5 nodelay;
proxy_connect_timeout 10s;
proxy_read_timeout 30s;
proxy_pass http://callback_backend;
}
This example assumes nginx is the public edge and should overwrite client-supplied forwarding headers. In a real proxy chain, configure trusted proxies and real-IP processing first; only then preserve the chain you actually trust.
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Forwarded-Host $host;
proxy_set_header Proxy "";
nosniff and a deliberate referrer policy are low-drama defaults. HSTS is powerful but sticky: enable it only when HTTPS is genuinely permanent for that host. CSP and framing policy are application-specific and should be designed, not copied from a random checklist.
add_header X-Content-Type-Options "nosniff" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
add_header Strict-Transport-Security "max-age=31536000" always;
# Example only if the application should never be framed:
add_header Content-Security-Policy "frame-ancestors 'none'" always;
JSON escaping, request IDs, protocol, status, latency and upstream timings make incident analysis much easier. Keep secrets, authorization material and full session state out of the default access format; use a separate protected debug path if you truly need deeper capture.
log_format edge_json escape=json
'{"time":"$time_iso8601",'
'"request_id":"$request_id",'
'"remote_addr":"$remote_addr",'
'"host":"$host",'
'"method":"$request_method",'
'"uri":"$uri",'
'"status":$status,'
'"request_time":"$request_time",'
'"upstream_status":"$upstream_status",'
'"upstream_time":"$upstream_response_time"}';
A TLS default server can reject unknown SNI before normal HTTP handling on TCP. HTTP/3 arrives on a separate QUIC listener, so give that socket its own default server and a cheap rejection path as well.
server {
listen 443 ssl default_server;
server_name _;
ssl_reject_handshake on;
}
server {
listen 443 quic reuseport default_server;
server_name _;
return 444;
}
Keep Alt-Svc host-scoped rather than global, validate client addresses with QUIC Retry when appropriate, and verify the UDP path independently from the existing TCP/H2 path.
quic_retry on;
map $host $h3_alt_svc {
default "";
public.example 'h3=":443"; ma=3600';
}
add_header Alt-Svc $h3_alt_svc always;
Some headers make sense everywhere; CSP, cache policy and indexing often do not. A host map lets nginx keep a shared edge while emitting the policy that actually matches each application.
map $host $csp_header {
default "";
public.example "default-src 'self'; object-src 'none'";
}
map $host $robots_header {
default "";
private.example "noindex, nofollow";
}
add_header Content-Security-Policy $csp_header always;
add_header X-Robots-Tag $robots_header always;
A compact all-request JSON log is useful for baselines. A second conditional log containing only 4xx/5xx requests gives incident tooling a smaller stream without changing enforcement or forcing verbose debug logging globally.
map $status $log_error_request {
default 0;
~^[45] 1;
}
access_log /var/log/nginx/access.json edge_json;
access_log /var/log/nginx/error.json edge_json if=$log_error_request;
You can deliberately reject HTTP/1.x at a controlled frontend, but that is a compatibility decision, not a baseline security requirement. Some legitimate clients, integrations and recovery paths still depend on HTTP/1.1.
nginx still documents its HTTP/3 module as experimental. Test the exact nginx build, TLS library, UDP path, client-IP semantics and fallback behavior before advertising QUIC.
map $server_protocol $legacy_http {
default 0;
"HTTP/1.0" 1;
"HTTP/1.1" 1;
}
server {
if ($legacy_http) {
return 505;
}
}
For request classes that outgrow map and geo, a deterministic access-phase dispatcher keeps policy modules explicit. Run it once per request, allowlist module names, validate module contracts and fail closed if the policy engine itself cannot be evaluated.
-- access.lua
if ngx.ctx.edge_policy_ran then
return
end
ngx.ctx.edge_policy_ran = true
local allowed = {
collect_client = true,
request_policy = true,
}
local function fail_closed(reason)
ngx.log(ngx.ERR, "edge_policy reason=", reason)
ngx.header["Cache-Control"] = "no-store"
return ngx.exit(ngx.HTTP_SERVICE_UNAVAILABLE)
end
for name in ngx.var.edge_chain:gmatch("([^,]+)") do
if not allowed[name] then
return fail_closed("module_not_allowed")
end
local ok, mod = pcall(require, name)
if not ok or type(mod.run) ~= "function" then
return fail_closed("module_contract_invalid")
end
local ran, handled = pcall(mod.run, {})
if not ran then
return fail_closed("module_runtime_failed")
end
if handled == true then
return
end
end
If the public edge prefers newer protocols but a crawler or service still needs HTTP/1.1, verify the thing that justifies the exception. A crawler claim can require both an expected UA and a published network; an export client can require the exact export route plus its expected client shape.
map "$verified_crawler:$is_legacy_service_route:$service_client_ok" $http1_exception {
default 0;
"1:0:0" 1;
"0:1:1" 1;
}
map "$legacy_http:$http1_exception" $reject_http1 {
default 0;
"1:0" 1;
}
Separate low-volume streams for verified crawlers, explicit allowlists, rate-limit events and rejected policy classes make rotation review much faster. Lua can enrich only the deep-debug class in the log phase, so ordinary traffic stays cheap.
log_by_lua_file /etc/nginx/lua/log_client.lua;
access_log /var/log/nginx/allowlist.json edge_json if=$log_allowlist;
access_log /var/log/nginx/goodbots.json edge_json if=$log_goodbot;
access_log /var/log/nginx/ratelimit.json edge_json if=$log_ratelimit;
access_log /var/log/nginx/policy.json edge_json if=$log_policy_event;
access_log /var/log/nginx/access.json edge_json if=$log_normal;
Once the obvious surface is finite, the next layer is disagreement: frontend versus upstream parsing, SNI versus Host, cache versus authentication state, DNS versus routing, retries versus side effects, and nginx versus the filesystem it is allowed to touch.
Do not pass client hop-by-hop headers or forwarding claims through unchanged. Let nginx parse the request, then construct a small, explicit upstream request. Parser disagreement is where desync classes become interesting.
TLS SNI chooses a certificate and often a virtual server; HTTP Host chooses application routing. When an application expects them to describe the same public service, mismatches deserve an explicit policy instead of accidental routing.
A cache key is part of access control whenever responses vary by cookie, authorization, host, language or another request property. Never cache personalized output just because the upstream forgot to say no.
If nginx resolves dynamic upstream names at runtime, the resolver becomes part of the trust boundary. Use a resolver you control or deliberately trust, bound resolution time and do not treat public DNS as a transparent implementation detail.
A GET retry is usually boring. A checkout, webhook or administrative POST may not be. Keep automatic upstream retry behavior conservative and never enable retries of non-idempotent methods without an application-level idempotency contract.
Large bodies and buffered upstream responses can spill to temporary files. Their directories, permissions, capacity and mount policy are part of edge hardening just as much as header limits are.
Symlinks are useful for deployments, but on writable or multi-tenant trees they can also cross trust boundaries. Decide deliberately whether a root may follow them; do not bolt on disable_symlinks without checking how the site is published.
The nginx worker should have the filesystem and kernel capabilities it actually needs — not a convenient view of every home directory, device, credential and writable path on the host.
Keep nginx and upstream software current, use one well-defined frontend path, strip the historical Proxy header, and set forwarding and hop-by-hop headers yourself. WebSocket upgrade headers belong only on the routes that actually upgrade. Turning off chunked transfer encoding globally is not a generic request-smuggling fix.
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Proxy "";
proxy_set_header Connection "";
# Only in an actual WebSocket location:
# proxy_set_header Upgrade $http_upgrade;
# proxy_set_header Connection "upgrade";
A fail-closed default server handles unknown SNI. For applications that require one canonical public name, a second policy can also reject a request whose TLS name and HTTP Host do not belong to the same allowed pair. Canonical redirect hosts should be modeled deliberately rather than accidentally caught by this check.
map "$ssl_server_name|$host" $name_pair_ok {
default 0;
"www.example.net|www.example.net" 1;
"example.net|example.net" 1;
}
if ($name_pair_ok = 0) {
return 421;
}
Keep header buffers finite and preserve strict parsing unless the application has a documented reason not to. nginx uses some request-line and header directives during virtual-server selection, so changing them per host can be less local than it looks.
client_header_buffer_size 2k;
large_client_header_buffers 4 16k;
client_header_timeout 10s;
ignore_invalid_headers on;
underscores_in_headers off;
Large static objects may need byte ranges for media players or resumable downloads. Most HTML/API responses do not need arbitrarily many ranges. Limit or disable them where the workload does not use multipart range responses.
# Most dynamic/API output:
max_ranges 1;
# Or disable byte ranges completely where they have no use:
# max_ranges 0;
If authenticated state, cookies, language or request headers can change the response, either include the relevant dimension in the cache design or do not cache that request. Keep the cache key readable and explicit; never let an untrusted Host or a hidden personalization input create shared output for another user.
map $http_authorization $has_authorization {
default 1;
"" 0;
}
proxy_cache_key "$scheme|$request_method|$host|$request_uri";
proxy_cache_bypass $has_authorization $cookie_session;
proxy_no_cache $has_authorization $cookie_session;
Runtime DNS resolution is useful for changing service addresses, but it moves routing decisions into DNS. Point nginx at a secured local or private resolver, use a bounded cache lifetime appropriate for the environment, and keep resolution timeouts short enough that DNS failure does not become request pile-up.
# 192.0.2.53 is documentation space; use your trusted resolver.
resolver 192.0.2.53 valid=60s;
resolver_timeout 2s;
Keep connect/read/send timeouts finite and be conservative with retries. nginx does not retry non-idempotent requests unless configured to do so; keep it that way unless the application has a real idempotency key or provider reconcile contract. “The socket failed” is not proof that the upstream did nothing.
proxy_connect_timeout 5s;
proxy_read_timeout 30s;
proxy_send_timeout 30s;
proxy_next_upstream error timeout;
proxy_next_upstream_tries 2;
# Do not add: non_idempotent
nginx writes bodies larger than its buffer to temporary files. Put those files on a deliberately sized path with restrictive ownership; on Linux, a dedicated filesystem with nodev, nosuid and usually noexec can reduce the consequences of an application mistake. Increase limits only on routes that really upload.
client_max_body_size 1m;
client_body_temp_path /var/lib/nginx/client-body 1 2;
location /upload/ {
client_max_body_size 64m;
client_body_timeout 30s;
# auth / method / content-type policy here
}
A very small list of paths that cannot exist on the service — exposed VCS metadata, environment files or a framework the application does not run — can be rejected at the edge. Keep this list short and high-confidence. It is a cost sink, not a replacement for application security.
map $uri $impossible_probe {
default 0;
~*^/(?:\.env|\.git)(?:/|$) 1;
~*^/(?:wp-login\.php|xmlrpc\.php)(?:/|$) 1;
}
if ($impossible_probe) {
return 444;
}
When nginx itself fronts password, subrequest, JWT or OIDC authentication, auth_delay can make 401 timing less informative. It is not password hashing and it is not rate limiting; it is one small way to avoid turning rejection latency into an oracle.
location /private/ {
auth_request /_auth;
auth_delay 200ms;
}
location = /_auth {
internal;
# loopback/private auth upstream
}
nginx already drops worker privileges, but the service manager can narrow the process view further. Treat these as deployment policy, not copy-paste defaults: certificate loading, dynamic modules, Unix sockets, log paths and temp directories all need to be included deliberately.
# systemctl edit nginx
[Service]
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictSUIDSGID=true
# Then explicitly grant only required writable paths.
# ReadWritePaths=/run/nginx /var/cache/nginx /var/lib/nginx /var/log/nginx
For applications with a non-trivial CSP, start new policy in Content-Security-Policy-Report-Only, inspect violations, then enforce. Treat the report endpoint itself as untrusted input: body limits, rate limits and retention still apply.
add_header Content-Security-Policy-Report-Only
"default-src 'self'; object-src 'none'; base-uri 'self'; report-to csp" always;
# Keep the reporting endpoint small and separately budgeted.
A useful public baseline does not need to reveal the production-specific detection layer, internal topology or the exact things an operator watches for.
The page now shows trust classes, CIDR verification, policy composition and Lua orchestration. What stays private are the exact abuse signatures, path/referrer combinations, state-token mechanics and bypass-sensitive matcher thresholds used by a specific production application.
The guide now shows the lane architecture and trust-boundary pattern because those ideas are educational. Real hostnames, internal addresses, certificate paths and deployment-specific routing details still add no value here.
Cookies, authorization headers, tokens and raw request bodies are not part of a public logging example. Debug capture needs its own access and retention policy.
User agents, IPs, TLS fingerprints and referrers are useful signals. They are not reliable identity and should not become a universal trust root.
The right policy still needs a safe rollout. Validate the effective config, stage compatibility-sensitive controls and watch the actual rejection classes after reload.
nginx -t should be boring. Use nginx -T privately when you need to inspect the effective include tree; do not publish that output blindly.
nginx supports dry-run modes for native request and connection limits. They are ideal for learning the real excess rate before choosing a rejection threshold.
Browser traffic is not the whole internet. Test APIs, webhooks, package clients, monitoring, bots you intentionally support and any old devices you actually care about.
Track 4xx/429 classes, upstream latency and connection behavior immediately after rollout. A sudden clean drop in traffic can be a successful block or a successful self-denial.
A few primary references behind the examples above. Version details matter, especially around QUIC/HTTP/3 and newer inheritance behavior.