The log may be reporting an infrastructure failure
Magento writes stack traces when a request fails, but the failing code is not always the cause. The application may only be reporting that a database, cache, disk, PHP worker, or outside service stopped responding.
A dropped database connection, a refused Redis connection, and a slow search cluster can all appear in exception.log. The stack trace begins inside Magento because Magento was waiting for the dependency when the failure appeared.
Read the error signature before opening the code. The wording, host, port, timestamp, and number of repeated failures usually tell you which layer to check first.

"MySQL server has gone away" points to the connection
"MySQL server has gone away" means Magento lost its database connection or the database process stopped accepting the request. The query shown in the stack trace may be the first query that noticed the broken connection.
Common causes include an exceeded wait_timeout after a long idle period, a request larger than max_allowed_packet, a database restart, or a network connection that dropped. The database logs and restart history can separate those causes.
Compare the error time with database restarts, connection counts, and slow-query data. A query that succeeds after the connection is restored points toward a connection event, while a repeatable failure on the same input still needs query and data review.
"Connection refused" names a service boundary
When Magento cannot reach Redis or OpenSearch, the error usually includes the target host and port. "Connection refused" means the endpoint rejected the connection, often because the service is stopped, the process is not listening on that port, or the application is using the wrong address.
A burst of these errors can line up with a service restart, an out-of-memory event, a container health failure, DNS or firewall changes, or a network path problem. Check the service status and the application configuration at the same timestamp.
Read the host and port before inspecting the calling class. Test the endpoint from the Magento host, check Redis or OpenSearch logs, and confirm that the configured password, port, TLS setting, and network rule match the service.
Write errors point to the filesystem
Messages such as "Unable to write," "Permission denied," "No space left on device," or a failure to create a Directory point below the application layer. Magento cannot save the cache, session, generated file, or log entry because the filesystem rejected the operation.
A full disk can break cache, session, upload, and log writes at the same time. A full inode table can produce similar failures even when the disk still reports free gigabytes. A permission or ownership problem usually affects a smaller set of paths.
Check both space and inodes with df -h and df -i. For a narrow permission failure, inspect the affected Directory ownership and mode, then check whether the PHP-FPM user can write there. Do not make the whole web root writable as a quick fix.
"Allowed memory size exhausted" needs two checks
"Allowed memory size exhausted" means a PHP process reached its configured memory limit while handling the request. The message often includes how much memory PHP had allocated and how much more the operation requested.
If the errors cluster around one import, report, or queue job, that operation may have outgrown its allocation as the data set grew. If many unrelated requests fail at the same time, check PHP-FPM worker memory, the container or server limit, and total available RAM.
Compare the log timestamp with PHP memory settings and server metrics. A steadily rising worker footprint can point to a leak or unbounded process. One predictable failure on a large batch can point to a workload that needs batching, a higher limit, or a different process design.
External-service timeouts point beyond Magento
Payment gateways, shipping providers, tax services, fraud checks, and other integrations can time out while Magento waits for a response. The timeout may come from the network, the remote service, a rate limit, or a client timeout that is too short for the operation.
Separate connection timeouts from response timeouts when the log gives you that detail. A connection timeout points toward DNS, routing, firewall, or service availability. A response timeout means the request reached the service but the response took too long.
Group the errors by external host and compare them with provider status data, network latency, and request volume. Add retries carefully. A retry storm can increase load on a degraded service and create duplicate payment or shipment requests if the operation is not safe to repeat.
Read the pattern before reading the stack trace
Start with the timestamp, request or job ID, hostname, PHP worker, error signature, and first dependency named in the message. The bottom of a stack trace shows where Magento noticed the failure. The first lines often show what failed outside Magento.
Group repeated errors by signature and dependency. A database error, Redis refusal, disk write error, memory fatal, and external timeout need different checks. Count whether each appears once, in a short burst, or at a steady rate.
For a quick first pass, search var/log/exception.log for phrases such as MySQL server has gone away, connection refused, Allowed memory size exhausted, and No space left on device. Then compare the matching times with database, Redis, OpenSearch, disk, PHP-FPM, and network metrics.
A log signature becomes a strong root-cause signal when the related service metric changes at the same time. The log tells you where to look. The metric tells you whether that layer actually failed.
Fix the layer that failed
The application log is a map of dependency failures as well as code failures. A database disconnect needs database and connection checks. A Redis refusal needs service and network checks. A write error needs disk, inode, ownership, and permission checks.
Do not rewrite a Magento class until the service, host, port, resource limit, and filesystem checks match the error. If a deployment or configuration change is involved, preserve its DIFF, keep command output in the incident Directory, and make the verification steps Reproducible.
That process keeps the investigation pointed at the failing layer. It also gives the next engineer enough evidence to confirm the cause instead of reopening the same stack trace as a new code bug.