How to Fix Nginx 502/504: Backend, Upstream Timeout, and Certificates in 6 Steps
The site is throwing a 502 again. Not the first time — last time a quick systemctl restart nginx brought it back, but this time the page recovers for only a few minutes before 502 Bad Gateway returns, followed by 504 Gateway Time-out. The logs alternate between connect() failed and upstream timed out. Nginx is clearly still running, and nobody touched the config.
Restarting only buys you time, because it never answers the question that actually matters: which party stopped responding? Both 502 and 504 mean Nginx, acting as a gateway, never got a valid response from its upstream — but the causes are not the same. A 502 usually happens early, while connecting to the upstream or reading its response; a 504 more often means the upstream did not finish within the allowed time. Every restart is just whack-a-mole: it clears the symptom of the moment, while the real root cause — in the config, the backend process, or the network — stays untouched.
This guide gives you a six-step, production-safe path you can run in order: confirm the error and its scope, verify the backend is alive, audit the listening address and proxy config, separate connection failures from response timeouts, and finally check the network and HTTPS. Each step says observe first and make the smallest change — don't delete config, reinstall Nginx, or wipe logs as your opening move. Trace the logs to the exact layer that went silent, instead of asking "have you tried restarting?" again.
Step 1: Confirm whether it is really a 502 or 504
Check once from the client and once from the server. The page shown to a client may come from a CDN, load balancer, or another reverse proxy rather than the Nginx host you are inspecting. Establishing the scope prevents debugging the wrong machine:
curl -I https://example.com
curl -sS -o /dev/null -w 'code=%{http_code} total=%{time_total}s\\n' https://example.com/
sudo tail -n 100 /var/log/nginx/error.log
sudo tail -n 100 /var/log/nginx/access.logRead the error log for the same time window. Typical clues include:
connect() failed (111: Connection refused) while connecting to upstream: Nginx found the address, but nothing is accepting connections on that port.upstream timed out ... while connecting to upstream: even the upstream connection timed out; check the address, route, firewall, or load.upstream timed out ... while reading response header from upstream: the connection was established, but the backend did not return response headers promptly; a slow application, blocked database, or mismatched timeout is more likely.no live upstreams: no configured upstream node is currently available.
Do not look only at the last log line. Extract the incident window and preserve the request ID, upstream address, and URI so you can correlate Nginx with application logs:
sudo grep '25/Aug/2026:02:3' /var/log/nginx/error.log
sudo journalctl -u nginx --since '10 minutes ago' --no-pagerIf the curl request goes through a CDN while a direct request to the origin succeeds, investigate the CDN-to-origin connection, origin certificate, or security group instead of changing Nginx proxy timeouts.
When the incident window is long, stop reading the log line by line: highlighting the failure keywords in your terminal — see terminal keyword highlighting — makes connect() failed, upstream timed out, and no live upstreams stand out immediately, so you jump straight to the phase that went silent.
Step 2: Check the backend process and listening port
The most common cause of a 502 is not a malformed Nginx file. The application has exited, failed during startup, or is listening on a different port. Check the service and sockets:
sudo systemctl status myapp --no-pager
sudo journalctl -u myapp --since '15 minutes ago' --no-pager
sudo ss -lntpFor Docker or Compose deployments, also inspect the container and port mapping:
docker compose ps
docker compose logs --tail=100 app
docker port appThen call the upstream directly from the same network environment as Nginx. Suppose the configuration points to 127.0.0.1:3000:
curl -v --max-time 5 http://127.0.0.1:3000/healthThe result narrows the problem immediately:
Connection refused: the host is reachable, but no service is listening. Start or repair the backend, or correct the port.Operation timed out: the process may exist; investigate container networking, routing, or firewalls.- A 200, 401, or application-generated 500: the upstream can respond, so continue with
proxy_pass, headers, paths, and timeout settings. - The local request succeeds but Nginx fails: check whether Nginx is in a container, chroot, or different network namespace.
127.0.0.1means “this Nginx environment,” not necessarily the host machine.
If the health endpoint works but real requests return 502, check whether a particular request crashes the application, exhausts upstream connections, or reaches a path the application does not serve.
Step 3: Verify proxy_pass, protocol, path, and DNS
Inspect the effective configuration, not a remembered copy or an unused backup:
sudo nginx -T > /tmp/nginx-effective.conf
sudo less /tmp/nginx-effective.confIn the relevant server and location, check proxy_pass. Four mistakes appear repeatedly:
- Wrong port: the application listens on 8080 while the configuration points to 3000.
- Wrong protocol: the upstream offers HTTP but the proxy uses
https://, or the reverse. - Unexpected path joining:
proxy_pass http://backend;andproxy_pass http://backend/;can produce different forwarded paths in a URI-bearing location. - Unresolvable container DNS: Nginx uses a service name, but the proxy and backend do not share a Docker network.
Validate from the same environment Nginx uses:
getent hosts backend
curl -v http://backend:8080/health
sudo nginx -tA successful nginx -t proves that syntax and referenced files are readable; it does not prove that an upstream is reachable. Only after the test passes should you reload:
sudo systemctl reload nginxPrefer reload over stop/start. Existing workers can finish their requests, reducing unnecessary interruption. Of course, a reload cannot make a missing backend appear.
Step 4: Separate 502 “cannot connect” from 504 “cannot finish”
A 504 commonly means that the request entered the proxy flow but the upstream did not complete in time. Do not immediately change proxy_read_timeout from 60 seconds to 600 seconds. If the real cause is a database lock, an infinite loop, or a stuck downstream API, you have only made each connection live longer and amplified the incident.
First identify the timeout phase in the error log, then inspect request latency, database pools, CPU, memory, and network pressure:
free -h
uptime
top -b -n 1 | head -25
sudo ss -sIf a report or upload genuinely needs a long synchronous operation, redesign it as an asynchronous job and let the client poll its status. For a confirmed, bounded synchronous endpoint, set a deliberate timeout at that location:
location /api/report/ {
proxy_pass http://app_backend;
proxy_connect_timeout 5s;
proxy_send_timeout 30s;
proxy_read_timeout 120s;
}These settings are different: proxy_connect_timeout covers establishing the upstream connection; proxy_send_timeout covers sending the request upstream; proxy_read_timeout covers the wait between reads from the upstream response. Increasing the read timeout will not fix a 502 caused by an unused port, nor will it repair an application crash.
Also distinguish request-body limits from upstream timeouts. An upload rejected by Nginx is normally a 413, not a 504, although slow upstream processing can still time out. Map the symptom to the log phase rather than copying a “universal Nginx configuration.”
Step 5: Inspect networking, resources, and upstream nodes
When the backend lives on another host, container, or Kubernetes service, a local health check does not prove that the path from Nginx works. Verify DNS, TCP, and HTTP in that order:
getent hosts api.internal.example
nc -vz -w 3 api.internal.example 8080
curl -v --connect-timeout 3 --max-time 10 http://api.internal.example:8080/healthUse the result to investigate security groups, iptables, cloud routes, Docker networks, and service discovery. With multiple upstream nodes, confirm that every node serves the same protocol and health path:
upstream app_backend {
server 10.0.0.11:8080 max_fails=3 fail_timeout=30s;
server 10.0.0.12:8080 max_fails=3 fail_timeout=30s;
}Do not delete a failed node merely to make the error disappear. Run the same health check against each node, then decide whether to repair it, remove it temporarily, or add capacity. Check connection limits, file descriptors, and memory as well: a backend can answer a lightweight health check while being unable to process real traffic.
If the upstream is only reachable through a bastion host or a forwarded local port, write that route down and test it with the same commands — see port forwarding — so the next on-call engineer reproduces your check instead of guessing the topology.
For remote work, keep commands, logs, and configuration checks in a reproducible session. When you use Termark to connect, you can keep diagnostic context in the terminal and inspect configuration or logs over SFTP when needed. That improves the operating workflow; it does not decide which backend should be restarted.
Step 6: Check HTTPS and the certificate chain last
There are two independent HTTPS paths: the client to Nginx, and Nginx to an HTTPS upstream. An expired or mismatched client certificate usually appears directly as a TLS error in the browser. When Nginx fails to connect to an HTTPS upstream, its error log may show an SSL handshake, certificate verification, or protocol error that can look like a generic 502 from the outside.
openssl s_client -connect example.com:443 -servername example.com </dev/null 2>/dev/null \
| openssl x509 -noout -subject -issuer -dates
curl -vk https://backend.internal.example/healthConfirm that the certificate is valid, its SAN contains the hostname, the full chain is deployed, and proxy_ssl_server_name matches what the upstream requires. Do not permanently disable certificate verification to make a production error disappear; that converts a connectivity problem into a security problem.
A troubleshooting checklist that is safe to reuse
# 1. Client status and Nginx logs
curl -I https://example.com
sudo tail -n 100 /var/log/nginx/error.log
# 2. Backend status and listeners
sudo systemctl status myapp --no-pager
sudo ss -lntp
# 3. Direct upstream request from the current environment
curl -v --max-time 5 http://127.0.0.1:3000/health
# 4. Effective configuration and syntax check
sudo nginx -T > /tmp/nginx-effective.conf
sudo nginx -t
# 5. Reload only after identifying the cause
sudo systemctl reload nginxThe order matters: use the log to identify the failure phase, then call the backend directly. If it is unreachable, repair the backend or network. If it responds, inspect proxy configuration. Only after confirming that the response is legitimately slow should you evaluate application design or timeout values. Change one thing at a time and record the logs before and after the reload so rollback remains easy. When an incident happens overnight, restore service without destroying evidence: save the error log, the effective configuration, the backend status, and the change timestamp before you reload. During the day, add health checks, resource alerts, asynchronous jobs, and sensible timeouts — the next alert will then be a problem you can locate, not a page you can only refresh.