~0x64616E69656C
project 04 · nginx edge-security field guide

Less magic.
Smaller surface.
Boring on purpose.

A practical field guide for internet-facing nginx: split protocol risk before application routing, reject what should never arrive, bound what can become expensive, keep trust explicit, use Lua as a small policy layer, and make the remaining behavior observable. No WAF theatre. No giant signature dump. No claim that one config fits every application.

Scope: edge architecture + common-sense hardening.
Examples are generalized from real operational patterns.
Exact production matchers, state tokens and bypass-sensitive rules stay private.
stream shunt · protocol lanes · phase machine · lua policy example policy · tune before production
01 / BASELINE

Start with invariants, not tricks.

The useful hardening decisions are usually boring: know which hosts, methods, protocols and request sizes you intend to serve, then make everything else explicit.

01

Fail closed

Unknown names should not fall through to a real application. Default servers exist to reject ambiguity, not to become accidental catch-alls.

02

Bound cost

Requests consume sockets, buffers, upstream time and application work. Put sensible ceilings around those resources before you need them.

03

Allow what exists

A static site does not need POST, CONNECT or TRACE. An upload endpoint does not need the same body limit as the rest of the site.

04

Trust deliberately

Forwarded addresses, scheme headers and proxy chains are security boundaries. Treat them as configuration, not ambient truth.

05

Observe cheaply

Request IDs, status, latency and upstream timing usually tell you more than indiscriminate request dumps — and are much safer to retain.

06

Stage policy

Dry-run limits where possible, inspect real traffic, then enforce. Hardening that breaks legitimate clients is just another outage.

02 / ARCHITECTURE

Hardening works best as a cost model.

The strongest controls are usually not giant blocklists. They make expensive work conditional, separate protocol risk from application policy, keep trust boundaries explicit and reject impossible traffic before it reaches application code.

01

Reject before routing

If a request is unambiguously impossible for the service, dispose of it before canonical redirects, proxying, sessions or application bootstrap consume work.

02

Budget by cost

A cheap static 404 and an expensive authenticated render do not need the same budget. Put the tighter control around the operation that actually burns CPU, DB time, sockets or third-party calls.

03

State beats fingerprints

IP addresses, user agents, TLS fingerprints and referrers are useful signals, not identity. When possible, bind enforcement to the costly resource, route or application state instead.

04

Count only what matters

nginx rate-limit zones can use an empty key to skip unrelated requests. That makes a narrow login, callback or expensive-route budget possible without taxing the whole site.

05

Detection is not trust

Telemetry may combine many weak signals to explain behavior. Enforcement should still use a small, reviewable contract and should not quietly promote a heuristic into an authentication decision.

06

Exceptions stay narrow

If a webhook, upload client or legacy protocol genuinely needs an exception, scope it to the exact host, route, method and budget instead of weakening the global edge.

07

QUIC is a separate edge

UDP/443 has its own listener and failure modes. Give unknown QUIC traffic a fail-closed default too; do not assume the TLS/TCP default server automatically covers it.

08

Internal auth stays internal

Authentication helpers and subrequest services do not need public reachability. Keep them on loopback or a private socket and expose only the application boundary that consumes their result.

09

Logs are a security boundary

Keep a compact default format for correlation, optionally split error-class requests into a second stream, and decide separately when deeper debug capture is justified.

Rule of thumb: detection can be broad and curious; enforcement should be narrow and boring. A useful signal does not automatically deserve to become a block key.
03 / PROTOCOL LANES

Do not make every protocol share the same risk budget.

One production pattern that has held up well is to split TCP/443 by ALPN before TLS termination: HTTP/2 enters a clean lane, HTTP/1.x enters a constrained compatibility cage, and HTTP/3 terminates separately on UDP/443 only where it is intentionally enabled.

TCP / 443

TLS ClientHello

The front listener only prereads enough TLS metadata to see ALPN. It does not terminate TLS here.

→
Stream shunt

ALPN dispatch

ssl_preread selects a loopback lane and PROXY protocol preserves the client address for the terminating server.

→
HTTP/2 lane

Clean path

TLS terminates here. Shared application policy still applies, but the transport does not inherit HTTP/1 compatibility penalties.

TCP / 443

No h2 ALPN

Legacy clients, control-plane tools and deliberate compatibility traffic still need a defined policy.

→
Stream shunt

Fallback dispatch

Traffic without the clean-lane ALPN is routed to a separate TLS listener instead of silently receiving the same policy.

→
HTTP/1 cage

Compatibility, constrained

Keepalive can be disabled, timeouts shortened, connection counts bounded and expensive routes given tighter budgets.

UDP / 443

QUIC Initial

HTTP/3 is a different transport edge. The TCP shunt cannot protect or route it.

→
Direct QUIC listener

Host-specific enablement

Only selected public application hosts advertise and terminate h3. Unknown QUIC traffic deserves its own fail-closed default.

→
HTTP/3 lane

Same app, distinct edge

Application authorization remains the same, while transport-specific limits, logging and rollout stay independently testable.

ALPN shunt

Split before TLS termination.

The stream layer can inspect ALPN from ClientHello without terminating TLS. That lets the HTTP servers remain normal TLS endpoints while still receiving distinct protocol classes and the original client address.

conceptual stream shunt
stream {
  map $ssl_preread_alpn_protocols $lane {
    ~\bh2\b   127.0.0.1:8443;
    default   127.0.0.1:9443;
  }

  server {
    listen 443 reuseport;
    ssl_preread on;
    proxy_protocol on;
    proxy_pass $lane;
  }
}
HTTP/1 cage

Treat H1 as a compatibility lane, not an equal peer.

A constrained H1 lane can keep old clients working without giving them the same socket lifetime and expensive-route budget as the preferred transport. The point is not “HTTP/1 is evil”; the point is that compatibility can have a smaller blast radius.

example cage policy
keepalive_timeout 0;
client_header_timeout 5s;
client_body_timeout 5s;
send_timeout 5s;
reset_timedout_connection on;

limit_conn h1_clients small_limit;

# Give expensive H1 routes their own tighter budget.
location ~ expensive_dynamic_route {
  limit_req zone=h1_expensive burst=small_burst nodelay;
  try_files $uri $uri/ /index.php?$args;
}
Shared include contract

Separate transport policy from application policy.

A useful file split is common-start → lane-specific policy → common-end. The common start restores the trusted client address and applies global request invariants; the lane file contains protocol-specific controls; the common end defines exact application endpoints, Lua access policy, static handling and logging.

lane composition
server {
  listen 8443 ssl proxy_protocol;
  http2 on;

  include edge/common_start.conf;
  include edge/h2_lane.conf;
  include edge/common_end.conf;
}

server {
  listen 9443 ssl proxy_protocol;

  include edge/common_start.conf;
  include edge/h1_cage.conf;
  include edge/common_end.conf;
}
Compatibility exceptions

Protocol support is not a protocol guarantee.

Browsers and modern crawlers may prefer h2 or h3, while inspection tools, provider control planes and legacy clients can still arrive over H1. Keep the compatibility exception conjunctive: protocol + verified network/client class + expected host + allowed method/route.

narrow H1 exception
map "$is_h1:$verified_crawler:$service_h1_ok" $allow_h1 {
  default 0;
  ~^1:1:   1;
  ~^1:0:1$ 1;
}

# The exception opens the protocol lane only.
# Method, body, host, rate and app authorization still apply.
Operational payoff: an H1 anomaly does not need to distort the H2/H3 policy. A QUIC experiment does not need to rewrite the TCP edge. And a legacy inspection client can receive one narrow compatibility exception without becoming globally trusted.
04 / TRUST CLASSES

Allowlisting without pretending it is identity.

Real edge policy usually needs more than one kind of exception. The useful pattern is to keep those classes separate: verified crawlers, fixed operational networks, route-bound service clients and weak telemetry signals should not collapse into one magic whitelist bit.

01

Goodbot is a policy class

A crawler can be treated as a low-risk exception signal because it arrives from a maintained provider network, because it is a narrowly recognized service client, or both. That can be enough to exempt it from a weak heuristic — but it is still not authentication.

02

Stronger exceptions need stronger evidence

When the exception itself is security-relevant — for example allowing an older protocol lane — combine the crawler claim with the matching provider CIDR. The same source can therefore be “good enough to skip a heuristic” yet still fail a stricter compatibility gate.

03

Service clients

Webhooks, inventory fetchers and preview clients often have recognizable request shapes. Treat that as an endpoint contract: exact route, method, content type, budget and application-level verification — not as a blanket trusted-agent rule.

04

Weak signals stay weak

ASN, country, referrer, Accept-Language, browser hints and TLS characteristics are useful for correlation. They may explain a request or contribute to a narrow gate, but they should not silently become authentication.

Crawler CIDRs

Pull network evidence from the provider, not from a blog post.

nginx geo maps are a convenient way to consume maintained CIDR files. In production I use separate generated files for public crawler networks such as Apple, Bing, DuckDuckGo, Google, OpenAI and Qwant, then validate the whole config before reload.

provider-maintained crawler networks
geo $crawler_net_ok {
  default 0;

  include /etc/nginx/cidrs/applebot.cidr;
  include /etc/nginx/cidrs/bingbot.cidr;
  include /etc/nginx/cidrs/duckduckbot.cidr;
  include /etc/nginx/cidrs/googlebot.cidr;
  include /etc/nginx/cidrs/google-special.cidr;
  include /etc/nginx/cidrs/openai-search.cidr;
  include /etc/nginx/cidrs/openai-user.cidr;
  include /etc/nginx/cidrs/qwantbot.cidr;
}
Claim + network

Require both when the exception depends on crawler identity.

A public CIDR says where the request came from. The user agent says which crawler role it claims. For stronger protocol or compatibility exceptions, require both. For lower-stakes heuristic exemptions, a separate broader goodbot signal can remain intentionally less strict.

verified crawler claim
map $http_user_agent $crawler_ua_claim {
  default 0;
  ~*Googlebot       1;
  ~*bingbot         1;
  ~*DuckDuckBot     1;
  ~*Applebot        1;
  ~*OAI-SearchBot   1;
  ~*ChatGPT-User    1;
}

map "$crawler_ua_claim:$crawler_net_ok" $verified_crawler {
  default 0;
  "1:1"   1;
}

# A broader low-risk signal can deliberately be OR-based.
map "$crawler_net_ok:$service_ua_ok" $goodbot_signal {
  default 0;
  ~^1:     1;
  ~:1$     1;
}
Purpose-bound whitelist

Keep operator, crawler and service exceptions independent.

Do not make one global $whitelist mean “skip security.” Use separate variables and combine them only where the endpoint actually needs the exception. Documentation-range addresses are used below deliberately.

separate trust classes
geo $ops_net {
  default 0;
  192.0.2.10/32 1;
  198.51.100.0/28 1;
}

map "$verified_crawler:$ops_net" $crawl_policy_exempt {
  default 0;
  ~^1:     1;
  ~:1$     1;
}

# A payment callback would use its own route-bound rule instead.
CIDR lifecycle

Automate the boring part, then fail closed on deployment.

Generated crawler lists should come from official feeds, be rendered into deterministic nginx include files, and be replaced per-file atomically. The updater should stop on fetch or parse errors and reload only after nginx -t succeeds. Staging the entire set as one transaction is an additional improvement when you need all feeds to advance together.

updater shape
#!/bin/sh
set -eu

tmpdir=$(mktemp -d)
trap 'rm -rf "$tmpdir"' EXIT

# fetch + validate provider JSON
# render: 203.0.113.0/24 1;

install -m 0644 "$tmpdir/provider.cidr" /etc/nginx/cidrs/provider.cidr.new
mv /etc/nginx/cidrs/provider.cidr.new /etc/nginx/cidrs/provider.cidr

nginx -t
service nginx reload
Important: “good bot” should normally mean “exempt from the specific heuristic that would misclassify it”, not “skip method checks, body limits, callback authentication, account authorization or application security”.
05 / LUA POLICY

Use Lua as a policy plane, not a shadow application.

The fast path should stay mostly native nginx: map, geo, exact locations and native limit zones. Lua earns its place when a small amount of stateful or structured policy is clearer there than in another fifty maps.

01

Classify

Cheap map/geo variables derive protocol, route, crawler, whitelist and cost classes.

02

Lane gate

H1/H2/H3 transport policy removes impossible traffic before application routing.

03

Dispatch

One access-phase dispatcher executes an explicit ordered module chain once per request.

04

Budget

Native request/connection zones count only the routes or resources that are actually expensive.

05

Application

Authentication, payment truth and business authorization remain application contracts.

06

Observe

Log-phase policy records the class and outcome without changing an already completed request.

01

Exactly once

Guard the dispatcher in ngx.ctx. A rewrite or internal redirect should not accidentally execute the access chain twice.

02

Explicit modules

Use an allowlist of known module names. Configuration should select from reviewed policy modules, not dynamically require arbitrary Lua paths.

03

Small contract

Each policy module gets one run(ctx)-style contract and a tiny set of valid return types. Ambiguous results are errors, not magic.

04

Fail closed

Missing modules, bad chains, duplicate modules and runtime exceptions should become a controlled edge failure — not a silent bypass.

05

Deep capture is selective

Do not parse every body and flatten every header on every request. Gate expensive diagnostic capture to the small class of requests that actually need it.

06

Log phase cannot become enforcement

Use the log phase for correlation and enrichment. Security decisions must already be complete before response logging begins.

Deterministic dispatcher

Make the Lua chain reviewable.

A compact dispatcher can normalize an ordered chain, reject duplicates, require an explicit module allowlist, verify the module contract and fail closed on load/runtime errors. That is much easier to reason about than scattered access_by_lua_block fragments.

generic access dispatcher
-- access.lua (simplified)
if ngx.ctx.edge_dispatched then return end
ngx.ctx.edge_dispatched = true

local allowed = {
  collect_context = true,
  route_policy    = true,
}

for name in configured_chain():gmatch("([^,]+)") do
  if not allowed[name] then
    return fail_closed("module_not_allowed", name)
  end

  local ok, mod = pcall(require, name)
  if not ok or type(mod.run) ~= "function" then
    return fail_closed("module_contract_invalid", name)
  end

  local ran, result = pcall(mod.run, ngx.ctx)
  if not ran then
    return fail_closed("module_runtime_failed", name)
  end
  if result == true then return end
end
Map first, Lua second

Keep cheap classification in nginx variables.

Large map/geo sets are lazy-evaluated by nginx. They are excellent for composing protocol, path, CIDR and method signals into a small number of policy variables. Lua should receive the already-classified request rather than reimplementing the whole edge parser.

native classification
geo $provider_net {
  default 0;
  include /etc/nginx/cidrs/provider.cidr;
}

map $request_uri $is_expensive_route {
  default 0;
  ~^/search 1;
}

map "$is_expensive_route:$provider_net" $access_chain {
  default      "collect_context,route_policy";
  "0:1"       "route_policy";
}
Selective deep telemetry

Debug the weird request, not the whole internet.

Detailed header/body capture is operationally useful, but expensive and sensitive. Derive a small $is_deep class from route/policy signals, capture only there, truncate bodies, and pass the minimum context through ngx.ctx to the log phase.

deep-capture gate
map "$is_asset:$verified_crawler:$interesting_request" $is_deep {
  default 0;
  "0:0:1" 1;
}

# access phase: normalize + truncate only when $is_deep = 1
# log phase: serialize the prepared context, no new security decision
Policy-aware logging

Log the reason class, not only the status code.

Separate streams for normal traffic, verified crawlers, rate limits, high-confidence sinks and deep diagnostics make investigations faster. A higher-confidence event can suppress duplicate lower-value log lines so one request does not masquerade as five incidents.

conditional fan-out
map $status $is_429 { default 0; 429 1; }
map $policy_sink $log_policy { default 0; ~.+ 1; }

access_log /var/log/nginx/access.json edge_json if=$log_normal;
access_log /var/log/nginx/policy.json edge_json if=$log_policy;
access_log /var/log/nginx/rate.json   edge_json if=$is_429;
Architectural rule: Lua may enrich, correlate and make narrow policy decisions. It should not quietly become a second routing framework, duplicate nginx-native limits, or turn heuristic telemetry into authentication.
06 / REQUEST LIFECYCLE

The file is not the program.

nginx is a phase machine. A readable edge design follows those phases deliberately: transport admission first, address normalization before policy inputs, native cost controls before application access, application truth after edge heuristics, and logging only after the outcome is known.

00 / STREAM

Transport admission

TCP/UDP, ALPN preread and protocol-lane selection happen before the HTTP request exists.

01 / POST-READ

Normalize address

Restore the trusted client address before downstream CIDR, rate and audit policy consumes it.

02 / REWRITE

Normalize request

Server/location rewrite logic can reject or transform early, before access and content work.

03 / PREACCESS

Bound cost

Native connection and request limiters execute here. The expensive request may die before access policy.

04 / ACCESS

Evaluate policy

Auth modules and access-phase Lua make narrow allow/deny decisions from already-normalized inputs.

05 / PRECONTENT

Resolve target

try_files and related routing decide whether static content, an application entry point or a miss is next.

06 / CONTENT

Application truth

Static output, FastCGI, proxying or a content handler executes. Payment/auth/business truth belongs here or deeper.

07 / LOG

Correlate outcome

Enrich and route telemetry after the response decision. Log phase should never retroactively become enforcement.

01

Directive order is not phase order

Two directives next to each other in a file may execute in different nginx phases. Review the module hook, not just the line number.

02

Normalize before classifying

Real client address, canonical path and protocol class should be stable before they become CIDR, limiter or logging inputs.

03

Spend budget before Lua

Use native preaccess limits for simple cost ceilings. Do not wake a richer policy engine merely to rediscover that a client exceeded a counter.

04

Exceptions grant one capability

An H1 compatibility exception grants H1. A crawler exception skips a weak crawler heuristic. Neither implies application authorization.

05

No trust escalation by telemetry

UA, CIDR, fingerprint and referrer can alter edge policy. They must not silently upgrade payment, account or session truth.

06

Later stages should narrow, not reopen

Once an impossible host, method, protocol or route is rejected, a later generic fallback should not resurrect it.

Phase-aware review

Review the effective request path, not the include tree.

A common split such as common-start → lane policy → common-end improves ownership, but the runtime still follows nginx phases. In a real edge, real-IP normalization can execute before a limiter even if its directive lives in another include; limit_req/limit_conn run in preaccess; access Lua runs later; try_files later again; log handlers run last.

mental execution model
# transport plane
accept -> preread -> lane

# HTTP plane
post-read     -> trusted client address
rewrite       -> early normalization / rejection
preaccess     -> native req/conn budgets
access        -> auth + Lua policy
precontent    -> try_files / route resolution
content       -> static / FastCGI / proxy
log           -> classify + correlate outcome

# config-file order is not a substitute for this model
Policy algebra

Compose capabilities instead of building a magic whitelist.

Keep lane admission, heuristic exemption, resource budgeting and application authorization as different outputs. This makes policy reviewable: one signal can relax one heuristic without mutating unrelated controls.

generic composition
map "$protocol:$verified_service" $lane_allowed {
  default 0;
  "h2:0" 1;
  "h3:0" 1;
  "h1:1" 1;
}

map "$verified_crawler:$route_class" $heuristic_exempt {
  default 0;
  "1:public" 1;
}

# application authorization is intentionally not derived here
07 / POLICY MATRIX

Different clients get different capabilities — not different truth.

A matrix makes exceptions auditable. The examples below are deliberately generic, but mirror a useful production principle: crawler, operator, callback and browser classes may receive different transport or heuristic treatment while method, body, resource and application contracts stay independently enforced.

Traffic classProtocol admissionEdge exemptionsCost budgetApplication truthTelemetry
Normal browser h2/h3 normalh1 cage None by default. Route/resource-specific native limits. Normal session, auth, CSRF and business rules. Normal access stream; deep capture only when selected.
Verified crawler h2/h3h1 only if compatibility policy allows May skip crawler-sensitive heuristics after provider verification. Still finite; crawler-specific budgets may differ. No account/payment authorization upgrade. crawler stream for volume and drift.
Operator / service allowlist Only the protocols the service actually needs. Narrow, named exemptions — never a universal bypass. Dedicated ceiling appropriate to the service. Service authentication and route authorization still required. allowlist stream with explicit reason class.
Provider callback Endpoint-bound compatibility, including legacy H1 only when required. Browser heuristics may be irrelevant on that exact endpoint. Dedicated method/body/rate envelope. provider signature/assert remains authoritative. Callback outcome + retry correlation; no secret material.
Internal auth subrequest not public internal location / loopback. Public crawler/whitelist policy does not apply. Small bounded internal budget. Auth service returns the authorization decision. Minimal auth result/latency; no credentials.
Unknown / weak H1 client cage or reject None. Tight connection, timeout and route budgets. Same application checks if it reaches content. Protocol/rejection class for compatibility analysis.
Repeated expensive resource Any otherwise-valid lane. No identity shortcut from IP/UA churn. state + canonical resource where appropriate. Business state remains unchanged by rate classification. 429 / cost stream with bounded context.
State-machine invariant: every edge transition should do one of three things: normalize context, narrow capabilities, or terminate the request. A later stage should not turn a weak signal into identity or silently re-enable a capability that an earlier stage removed.
08 / PATTERNS

A small, explainable edge.

These snippets are examples, not a drop-in complete config. Values are intentionally readable starting points; application requirements win.

Global request policy

Keep the default envelope finite.

Hide the nginx version, bound headers and bodies, and set finite read/write timeouts. Large uploads or long-running upstreams should get narrowly scoped exceptions instead of globally relaxed defaults.

http context
server_tokens off;

client_max_body_size 1m;
client_body_buffer_size 128k;
client_header_buffer_size 2k;
large_client_header_buffers 4 16k;

client_header_timeout 10s;
client_body_timeout 15s;
send_timeout 30s;
keepalive_timeout 30s;
TLS policy

Prefer current protocols; make replay a conscious choice.

TLS 1.2 + 1.3 remains a sensible compatibility baseline. Keeping 0-RTT disabled is conservative unless the application and every relevant request path are designed around replay semantics.

http or server context
ssl_protocols TLSv1.2 TLSv1.3;
ssl_session_cache shared:SSL:20m;
ssl_session_timeout 1d;
ssl_session_tickets off;
ssl_early_data off;
Unknown hosts

Do not serve an application by accident.

A dedicated TLS default server can reject handshakes for names you do not host. This is cleaner than letting an unknown SNI select whichever application server happens to be first.

fail-closed default vhost
server {
  listen 443 ssl default_server;
  server_name _;

  ssl_reject_handshake on;
}
Method surface

Static means GET and HEAD.

For a truly static vhost, other methods are noise. The return-only conditional below is intentionally simple; dynamic applications should define their method policy per endpoint instead.

map + static server
map $request_method $static_method_denied {
  default 1;
  GET     0;
  HEAD    0;
}

server {
  # A return-only if is easy to reason about.
  if ($static_method_denied) {
    return 405;
  }

  location / {
    try_files $uri $uri/ =404;
  }
}
Hidden files

Do not make dotfiles discoverable by default.

Reject hidden paths unless you deliberately publish one. If you use ACME HTTP-01 or a well-known metadata path, carve that path out explicitly rather than weakening the general rule.

server context
location ~ /\. {
  return 404;
}
Native budgets

Rate-limit work, not identities.

nginx has native request and connection budgets. Start with broad, measurable ceilings and narrow them only where the workload justifies it. If nginx sits behind another proxy, establish the trusted real-client address first or an IP-keyed limit is meaningless.

For HTTP/2 and HTTP/3, remember that nginx counts each concurrent request as a connection for limit_conn.

http + server context
limit_conn_zone $binary_remote_addr zone=per_ip_conn:10m;
limit_req_zone  $binary_remote_addr zone=public_req:10m rate=20r/s;

limit_conn_status 429;
limit_req_status 429;

server {
  limit_conn per_ip_conn 32;
  limit_req zone=public_req burst=60 nodelay;
}
Selective accounting

Use empty keys to limit one operation, not the whole site.

A login or sign-up budget should normally count only those requests. Build the key from the request shape you care about and return an empty value for everything else. nginx does not account empty keys in the zone.

targeted request budget
map "$request_method:$uri" $login_limit_key {
  default "";
  ~^POST:/user/(?:login|sign-up)$ $binary_remote_addr;
}

limit_req_zone $login_limit_key zone=login:1m rate=10r/m;

server {
  limit_req zone=login burst=10 nodelay;
}
Cost-aware state

Key expensive work to the resource being abused.

Per-IP controls are useful coarse ceilings, but they are often the wrong identity model. For stateful applications, a tighter budget can follow a server-issued session plus a canonical resource path so rotating egress addresses do not reset the cost accounting.

The exact key is application policy. Keep it out of generic public examples if it would reveal bypass-sensitive internals.

conceptual key design
# Coarse edge ceiling:
client address -> broad connection/request budget

# Narrow expensive-operation ceiling:
server-issued state + canonical resource
  -> cost budget for that exact operation

# Do not use as identity by itself:
UA / ASN / referrer / TLS fingerprint
Internal auth boundary

Do not publish an authentication helper just because nginx can proxy it.

Subrequest-based authentication works well when the helper is reachable only from nginx. Keep the auth location internal and the service on loopback or a private Unix socket; the public application consumes only the allow/deny result.

internal auth subrequest
location = /_auth {
  internal;
  proxy_pass http://unix:/run/app-auth.sock;
  proxy_pass_request_body off;
  proxy_set_header Content-Length "";
  proxy_set_header X-Original-URI $request_uri;
}

location /private/ {
  auth_request /_auth;
  proxy_pass http://app_backend;
}
Narrow compatibility exceptions

Make the weird client the exception, not the global policy.

Payment providers, webhooks, monitoring agents and old package clients sometimes need a protocol, method or timeout that normal browser traffic does not. Keep the exception on the exact route and preserve the stricter default everywhere else.

route-scoped exception
location = /provider/callback {
  # Only the provider contract gets the exception.
  limit_req zone=callbacks burst=5 nodelay;
  proxy_connect_timeout 10s;
  proxy_read_timeout 30s;
  proxy_pass http://callback_backend;
}
Reverse proxy trust

Do not forward client claims as facts.

This example assumes nginx is the public edge and should overwrite client-supplied forwarding headers. In a real proxy chain, configure trusted proxies and real-IP processing first; only then preserve the chain you actually trust.

proxy location
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Forwarded-Host $host;
proxy_set_header Proxy "";
Response headers

Use browser policy that matches the application.

nosniff and a deliberate referrer policy are low-drama defaults. HSTS is powerful but sticky: enable it only when HTTPS is genuinely permanent for that host. CSP and framing policy are application-specific and should be designed, not copied from a random checklist.

baseline headers
add_header X-Content-Type-Options "nosniff" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
add_header Strict-Transport-Security "max-age=31536000" always;

# Example only if the application should never be framed:
add_header Content-Security-Policy "frame-ancestors 'none'" always;
Observability

Log enough to explain a request.

JSON escaping, request IDs, protocol, status, latency and upstream timings make incident analysis much easier. Keep secrets, authorization material and full session state out of the default access format; use a separate protected debug path if you truly need deeper capture.

structured access log
log_format edge_json escape=json
  '{"time":"$time_iso8601",'
  '"request_id":"$request_id",'
  '"remote_addr":"$remote_addr",'
  '"host":"$host",'
  '"method":"$request_method",'
  '"uri":"$uri",'
  '"status":$status,'
  '"request_time":"$request_time",'
  '"upstream_status":"$upstream_status",'
  '"upstream_time":"$upstream_response_time"}';
QUIC default boundary

Fail closed on UDP/443 as deliberately as on TCP/443.

A TLS default server can reject unknown SNI before normal HTTP handling on TCP. HTTP/3 arrives on a separate QUIC listener, so give that socket its own default server and a cheap rejection path as well.

separate TCP and QUIC defaults
server {
  listen 443 ssl default_server;
  server_name _;
  ssl_reject_handshake on;
}

server {
  listen 443 quic reuseport default_server;
  server_name _;
  return 444;
}
HTTP/3 advertisement

Advertise QUIC only where you actually operate it.

Keep Alt-Svc host-scoped rather than global, validate client addresses with QUIC Retry when appropriate, and verify the UDP path independently from the existing TCP/H2 path.

host-scoped h3 policy
quic_retry on;

map $host $h3_alt_svc {
  default "";
  public.example 'h3=":443"; ma=3600';
}

add_header Alt-Svc $h3_alt_svc always;
Host-specific browser policy

Security headers are application policy, not one global string.

Some headers make sense everywhere; CSP, cache policy and indexing often do not. A host map lets nginx keep a shared edge while emitting the policy that actually matches each application.

per-host response policy
map $host $csp_header {
  default "";
  public.example "default-src 'self'; object-src 'none'";
}

map $host $robots_header {
  default "";
  private.example "noindex, nofollow";
}

add_header Content-Security-Policy $csp_header always;
add_header X-Robots-Tag $robots_header always;
Split observability

Keep a normal request stream and a cheap error mirror.

A compact all-request JSON log is useful for baselines. A second conditional log containing only 4xx/5xx requests gives incident tooling a smaller stream without changing enforcement or forcing verbose debug logging globally.

conditional error-class log
map $status $log_error_request {
  default 0;
  ~^[45]  1;
}

access_log /var/log/nginx/access.json edge_json;
access_log /var/log/nginx/error.json edge_json if=$log_error_request;
Optional compatibility policy

HTTP/2 + HTTP/3 only is not a universal hardening rule.

You can deliberately reject HTTP/1.x at a controlled frontend, but that is a compatibility decision, not a baseline security requirement. Some legitimate clients, integrations and recovery paths still depend on HTTP/1.1.

nginx still documents its HTTP/3 module as experimental. Test the exact nginx build, TLS library, UDP path, client-IP semantics and fallback behavior before advertising QUIC.

only when the compatibility envelope is known
map $server_protocol $legacy_http {
  default    0;
  "HTTP/1.0" 1;
  "HTTP/1.1" 1;
}

server {
  if ($legacy_http) {
    return 505;
  }
}
Lua access dispatcher

Use Lua as a small policy layer, not an unbounded second application.

For request classes that outgrow map and geo, a deterministic access-phase dispatcher keeps policy modules explicit. Run it once per request, allowlist module names, validate module contracts and fail closed if the policy engine itself cannot be evaluated.

generic access-phase dispatcher
-- access.lua
if ngx.ctx.edge_policy_ran then
  return
end
ngx.ctx.edge_policy_ran = true

local allowed = {
  collect_client = true,
  request_policy = true,
}

local function fail_closed(reason)
  ngx.log(ngx.ERR, "edge_policy reason=", reason)
  ngx.header["Cache-Control"] = "no-store"
  return ngx.exit(ngx.HTTP_SERVICE_UNAVAILABLE)
end

for name in ngx.var.edge_chain:gmatch("([^,]+)") do
  if not allowed[name] then
    return fail_closed("module_not_allowed")
  end

  local ok, mod = pcall(require, name)
  if not ok or type(mod.run) ~= "function" then
    return fail_closed("module_contract_invalid")
  end

  local ran, handled = pcall(mod.run, {})
  if not ran then
    return fail_closed("module_runtime_failed")
  end
  if handled == true then
    return
  end
end
Protocol exception matrix

Compatibility exceptions should be conjunctive, not hopeful.

If the public edge prefers newer protocols but a crawler or service still needs HTTP/1.1, verify the thing that justifies the exception. A crawler claim can require both an expected UA and a published network; an export client can require the exact export route plus its expected client shape.

narrow HTTP/1.1 exception
map "$verified_crawler:$is_legacy_service_route:$service_client_ok" $http1_exception {
  default 0;
  "1:0:0" 1;
  "0:1:1" 1;
}

map "$legacy_http:$http1_exception" $reject_http1 {
  default 0;
  "1:0"   1;
}
Conditional security streams

Log the policy class that fired, not only the final status.

Separate low-volume streams for verified crawlers, explicit allowlists, rate-limit events and rejected policy classes make rotation review much faster. Lua can enrich only the deep-debug class in the log phase, so ordinary traffic stays cheap.

policy-aware logging
log_by_lua_file /etc/nginx/lua/log_client.lua;

access_log /var/log/nginx/allowlist.json edge_json if=$log_allowlist;
access_log /var/log/nginx/goodbots.json edge_json if=$log_goodbot;
access_log /var/log/nginx/ratelimit.json edge_json if=$log_ratelimit;
access_log /var/log/nginx/policy.json edge_json if=$log_policy_event;
access_log /var/log/nginx/access.json edge_json if=$log_normal;
Do not cargo-cult the numbers. A 1 MiB body ceiling is great for a static site and wrong for a package registry. A 20 r/s budget may be generous for login and absurdly small for a tiled asset endpoint. The invariant is the hardening; the number is workload policy.
09 / DEEP EDGE

The awkward boundaries are where bugs survive.

Once the obvious surface is finite, the next layer is disagreement: frontend versus upstream parsing, SNI versus Host, cache versus authentication state, DNS versus routing, retries versus side effects, and nginx versus the filesystem it is allowed to touch.

01

One parser boundary

Do not pass client hop-by-hop headers or forwarding claims through unchanged. Let nginx parse the request, then construct a small, explicit upstream request. Parser disagreement is where desync classes become interesting.

02

Name the same service twice

TLS SNI chooses a certificate and often a virtual server; HTTP Host chooses application routing. When an application expects them to describe the same public service, mismatches deserve an explicit policy instead of accidental routing.

03

Cache is an auth boundary

A cache key is part of access control whenever responses vary by cookie, authorization, host, language or another request property. Never cache personalized output just because the upstream forgot to say no.

04

DNS is routing policy

If nginx resolves dynamic upstream names at runtime, the resolver becomes part of the trust boundary. Use a resolver you control or deliberately trust, bound resolution time and do not treat public DNS as a transparent implementation detail.

05

Retries can be mutations

A GET retry is usually boring. A checkout, webhook or administrative POST may not be. Keep automatic upstream retry behavior conservative and never enable retries of non-idempotent methods without an application-level idempotency contract.

06

Disk is part of request handling

Large bodies and buffered upstream responses can spill to temporary files. Their directories, permissions, capacity and mount policy are part of edge hardening just as much as header limits are.

07

Static roots need a symlink policy

Symlinks are useful for deployments, but on writable or multi-tenant trees they can also cross trust boundaries. Decide deliberately whether a root may follow them; do not bolt on disable_symlinks without checking how the site is published.

08

The worker is not the machine

The nginx worker should have the filesystem and kernel capabilities it actually needs — not a convenient view of every home directory, device, credential and writable path on the host.

Desync / request smuggling boundary

Reconstruct the upstream request instead of forwarding ambiguity.

Keep nginx and upstream software current, use one well-defined frontend path, strip the historical Proxy header, and set forwarding and hop-by-hop headers yourself. WebSocket upgrade headers belong only on the routes that actually upgrade. Turning off chunked transfer encoding globally is not a generic request-smuggling fix.

explicit proxy boundary
proxy_http_version 1.1;

proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Proxy "";
proxy_set_header Connection "";

# Only in an actual WebSocket location:
# proxy_set_header Upgrade $http_upgrade;
# proxy_set_header Connection "upgrade";
SNI / Host coherence

Do not let two routing names disagree silently.

A fail-closed default server handles unknown SNI. For applications that require one canonical public name, a second policy can also reject a request whose TLS name and HTTP Host do not belong to the same allowed pair. Canonical redirect hosts should be modeled deliberately rather than accidentally caught by this check.

explicit allowed name pairs
map "$ssl_server_name|$host" $name_pair_ok {
  default 0;
  "www.example.net|www.example.net" 1;
  "example.net|example.net"         1;
}

if ($name_pair_ok = 0) {
  return 421;
}
Header parser surface

Oversized and invalid headers should fail at the edge.

Keep header buffers finite and preserve strict parsing unless the application has a documented reason not to. nginx uses some request-line and header directives during virtual-server selection, so changing them per host can be less local than it looks.

finite header envelope
client_header_buffer_size 2k;
large_client_header_buffers 4 16k;
client_header_timeout 10s;

ignore_invalid_headers on;
underscores_in_headers off;
Range-request surface

Multipart ranges are a feature, not a birthright.

Large static objects may need byte ranges for media players or resumable downloads. Most HTML/API responses do not need arbitrarily many ranges. Limit or disable them where the workload does not use multipart range responses.

route-specific range policy
# Most dynamic/API output:
max_ranges 1;

# Or disable byte ranges completely where they have no use:
# max_ranges 0;
Cache poisoning boundary

Key every representation dimension — or bypass the cache.

If authenticated state, cookies, language or request headers can change the response, either include the relevant dimension in the cache design or do not cache that request. Keep the cache key readable and explicit; never let an untrusted Host or a hidden personalization input create shared output for another user.

conservative authenticated-cache bypass
map $http_authorization $has_authorization {
  default 1;
  ""      0;
}

proxy_cache_key "$scheme|$request_method|$host|$request_uri";
proxy_cache_bypass $has_authorization $cookie_session;
proxy_no_cache     $has_authorization $cookie_session;
Resolver trust

Dynamic upstream DNS needs an explicit trust anchor.

Runtime DNS resolution is useful for changing service addresses, but it moves routing decisions into DNS. Point nginx at a secured local or private resolver, use a bounded cache lifetime appropriate for the environment, and keep resolution timeouts short enough that DNS failure does not become request pile-up.

example trusted resolver
# 192.0.2.53 is documentation space; use your trusted resolver.
resolver 192.0.2.53 valid=60s;
resolver_timeout 2s;
Non-idempotent upstreams

Transport uncertainty must not become a double side effect.

Keep connect/read/send timeouts finite and be conservative with retries. nginx does not retry non-idempotent requests unless configured to do so; keep it that way unless the application has a real idempotency key or provider reconcile contract. “The socket failed” is not proof that the upstream did nothing.

bounded, idempotent retry policy
proxy_connect_timeout 5s;
proxy_read_timeout    30s;
proxy_send_timeout    30s;

proxy_next_upstream error timeout;
proxy_next_upstream_tries 2;

# Do not add: non_idempotent
Request-body temp files

Large uploads need a disk contract, not merely a bigger number.

nginx writes bodies larger than its buffer to temporary files. Put those files on a deliberately sized path with restrictive ownership; on Linux, a dedicated filesystem with nodev, nosuid and usually noexec can reduce the consequences of an application mistake. Increase limits only on routes that really upload.

small global body, explicit upload exception
client_max_body_size 1m;
client_body_temp_path /var/lib/nginx/client-body 1 2;

location /upload/ {
  client_max_body_size 64m;
  client_body_timeout 30s;
  # auth / method / content-type policy here
}
High-confidence pre-routing sinks

Reject impossible probes before redirects and upstream bootstrap.

A very small list of paths that cannot exist on the service — exposed VCS metadata, environment files or a framework the application does not run — can be rejected at the edge. Keep this list short and high-confidence. It is a cost sink, not a replacement for application security.

short, boring sink
map $uri $impossible_probe {
  default 0;
  ~*^/(?:\.env|\.git)(?:/|$) 1;
  ~*^/(?:wp-login\.php|xmlrpc\.php)(?:/|$) 1;
}

if ($impossible_probe) {
  return 444;
}
Authentication timing

Uniform denial can be more useful than a faster denial.

When nginx itself fronts password, subrequest, JWT or OIDC authentication, auth_delay can make 401 timing less informative. It is not password hashing and it is not rate limiting; it is one small way to avoid turning rejection latency into an oracle.

optional auth-response delay
location /private/ {
  auth_request /_auth;
  auth_delay 200ms;
}

location = /_auth {
  internal;
  # loopback/private auth upstream
}
Systemd / process sandbox

Constrain the service after the config is correct.

nginx already drops worker privileges, but the service manager can narrow the process view further. Treat these as deployment policy, not copy-paste defaults: certificate loading, dynamic modules, Unix sockets, log paths and temp directories all need to be included deliberately.

illustrative systemd override
# systemctl edit nginx
[Service]
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictSUIDSGID=true

# Then explicitly grant only required writable paths.
# ReadWritePaths=/run/nginx /var/cache/nginx /var/lib/nginx /var/log/nginx
CSP rollout telemetry

Stage browser policy before turning it into an outage.

For applications with a non-trivial CSP, start new policy in Content-Security-Policy-Report-Only, inspect violations, then enforce. Treat the report endpoint itself as untrusted input: body limits, rate limits and retention still apply.

report-only before enforce
add_header Content-Security-Policy-Report-Only
  "default-src 'self'; object-src 'none'; base-uri 'self'; report-to csp" always;

# Keep the reporting endpoint small and separately budgeted.
One more rule of thumb: the edge should make ambiguity smaller. If a setting creates a second interpretation of identity, routing, cache state or request framing, it deserves more scrutiny than a scanner rule ever will.
10 / BOUNDARY

What this page deliberately does not publish.

A useful public baseline does not need to reveal the production-specific detection layer, internal topology or the exact things an operator watches for.

Exact production matchers

The page now shows trust classes, CIDR verification, policy composition and Lua orchestration. What stays private are the exact abuse signatures, path/referrer combinations, state-token mechanics and bypass-sensitive matcher thresholds used by a specific production application.

Private topology

The guide now shows the lane architecture and trust-boundary pattern because those ideas are educational. Real hostnames, internal addresses, certificate paths and deployment-specific routing details still add no value here.

Secret-bearing logging

Cookies, authorization headers, tokens and raw request bodies are not part of a public logging example. Debug capture needs its own access and retention policy.

Identity-by-fingerprint

User agents, IPs, TLS fingerprints and referrers are useful signals. They are not reliable identity and should not become a universal trust root.

11 / DEPLOY

Hardening is an operations change.

The right policy still needs a safe rollout. Validate the effective config, stage compatibility-sensitive controls and watch the actual rejection classes after reload.

01 / VALIDATE

Test before reload.

nginx -t should be boring. Use nginx -T privately when you need to inspect the effective include tree; do not publish that output blindly.

02 / DRY RUN

Measure before enforcement.

nginx supports dry-run modes for native request and connection limits. They are ideal for learning the real excess rate before choosing a rejection threshold.

03 / COMPATIBILITY

Exercise real clients.

Browser traffic is not the whole internet. Test APIs, webhooks, package clients, monitoring, bots you intentionally support and any old devices you actually care about.

04 / OBSERVE

Watch what changed.

Track 4xx/429 classes, upstream latency and connection behavior immediately after rollout. A sudden clean drop in traffic can be a successful block or a successful self-denial.