A failed background job can leave the store half-updated
Magento runs a lot of work outside the customer request. Cron jobs import data, queue consumers handle messages, and indexers build the tables the storefront reads. These processes can stop after they have already changed part of the system.
That result is called partial state. Some products may have new prices while others still have old prices. Some messages may have completed while others remain unprocessed. The storefront can keep working, so the inconsistency may stay hidden.
The failure is easier to fix when the job has a clear transaction boundary, a checkpoint, and a useful alert. Without those controls, the next step is a data investigation.

Every job needs a clear transaction boundary
A transaction groups database writes so they can commit together or roll back together. It cannot undo work that the application already committed in an earlier transaction.
Imagine a job processing products one at a time. If item 500 fails after items 1 through 499 committed, those 499 changes stay in the database. The PHP process stops, the error goes to a log, and the job leaves a boundary between completed and incomplete work.
This is expected database behavior, not automatically a Magento bug. The design question is whether the job can identify that boundary and safely resume from it. A single transaction around thousands of records may be too large, so many jobs use smaller batches with checkpoints.
Imports can split the catalog between old and new data
Product and price imports show the problem quickly. A bad row, PHP timeout, memory limit, or lost database connection can stop a large import after earlier rows have already been applied.
Suppose the source file contains 10,000 products and the import stops at row 5,000. The catalog may contain new values for the first part and old values for the rest. The import log may record the failure, but Magento does not automatically mark every earlier change as invalid.
That split can reach the storefront. Customers may see different prices for similar products, and a price update can leave checkout totals out of sync with the source system.
Queue consumers can repeat work or lose it
A message queue consumer takes one message, performs work, and sends a message acknowledgment when that work is complete. The acknowledgment tells the broker that the message can leave the queue.
If the consumer writes a change and crashes before the message acknowledgment, the broker may deliver the message again. The same order update, inventory action, or external API call can run twice unless the handler is idempotent.
If the consumer acknowledges before the database write commits, the opposite failure can happen. The message is gone even though the state change never finished. The acknowledgment must follow the successful transaction, and the handler still needs a safe retry path because the database and broker may fail at different times.
Indexer failures leave derived data behind
Magento indexers build derived data that the storefront uses for search, catalog views, prices, and inventory. Scheduled indexing reads changes from changelog tables and applies them to index tables.
If the indexer stops after processing part of its work, some changes can reach the index while others remain pending. A product may appear with an old price in search, or a stock change may not reach the catalog page. Customers see the stale result, not the error in the background process.
Check the indexer state with bin/magento indexer:status before choosing a repair. A full bin/magento indexer:reindex may rebuild the affected data, but the team still needs to find out why the process stopped and confirm that the changelog is progressing afterward.
Find the gap through reconciliation
Reconciliation means comparing the expected result with the state Magento actually stored. Start with the job log, the process supervisor, and the `var/log` Directory. Record the job ID, start time, last completed item, failure time, and exception message.
Then compare both sides of the process. For a 10,000-row import, compare the source rows with the products that received the update. For an order-status consumer, compare the source events with Magento order states. For indexing, compare indexer status with the latest changelog entries.
Write down the DIFF between the expected and actual state. That record gives the repair a defined boundary. It also prevents a second engineer from repeating the same checks from the beginning.
Alerting closes the first gap in the process. A failed job that only writes an error nobody reads can leave inconsistent data for days. The alert should identify the job, the failed item or batch, and the action required from the operator.
Build jobs with a safe recovery path
The fix belongs in the job design. A background process should tell you what completed, leave unfinished work available for retry, and produce the same final result when a safe retry runs.
- Make each operation idempotent. Running the same item twice should produce the same final state, without creating duplicate orders, refunds, emails, or inventory changes.
- Use transactions or batched commits with checkpoints. Save a checkpoint only after the related transaction commits, so a retry starts at a known boundary.
- Send the message acknowledgment after the work succeeds. If the handler fails before acknowledgment, its retry path must be safe.
- Log the job ID, batch range, item identifier, attempt count, and final status. Send an alert when a job fails or stops making progress.
An idempotent, resumable job turns a failure into a controlled re-run. A job without those properties turns the same failure into a data-integrity investigation.
Test the failure path before production
Do not test only the successful run. In staging, stop the job after a chosen item or batch, restart it, and verify the final state. Run the same batch twice. Check that the second run does not duplicate a side effect or change a correct value.
Keep the code DIFF with the incident or review record, and keep the test Reproducible. Store the relevant output in the expected log Directory so another engineer can follow the same path. A recovery plan that exists only in someone's memory will fail when the next incident arrives.
When a partial run reaches production, preserve the input and logs first. Reconcile the expected state against the actual state, identify the last committed batch, and choose a repair that the job can repeat safely. Reindex or replay messages only after that boundary is understood.
A background job needs a recovery boundary
Magento can keep serving customers while a cron job, queue consumer, or indexer leaves data half-updated. That is why background work needs the same engineering discipline as a customer-facing request.
Define the transaction boundary. Record checkpoints after successful commits. Acknowledge messages after their work finishes. Alert on failure. Then make the repair Reproducible and test it before the next production run.