Dev.to Security 🔐 Cybersecurity 👁 0 📖 5 min read

When Skipping an API Error Can Delete Valid Data: A Cartography Fix

TL;DR: I fixed a BigQuery sync failure in Cartography, then narrowed the error tolerance after review showed that skipping failed enumeration could feed false deletions into graph cleanup. One malformed BigQuery table c

TL;DR: I fixed a BigQuery sync failure in Cartography, then narrowed the error tolerance after review showed that skipping failed enumeration could feed false deletions into graph cleanup.

One malformed BigQuery table could stop a Cartography GCP sync. My first patch let more API errors be skipped, but an automated reviewer caught a risk I had missed: a failed list request could leave cleanup deleting valid inventory nodes. I traced that concern through the sync code and changed which errors the patch would tolerate.

Cartography is a CNCF sandbox project written in Python that collects infrastructure assets and their relationships into a Neo4j graph. Its BigQuery ingestion lists resources, fetches extra details, loads the results, and runs cleanup. I worked on PR #2972 to fix the sync failure without making incomplete collection look like resource removal.

Why a missing graph node matters

Cartography's README describes queries for datastore access, network paths, and internet-exposed compute instances. The graph supplies the inventory and relationships behind those security questions. If cleanup removes a valid table node because enumeration failed, that table disappears from the inventory available to queries that depend on it.

The cloud resource has not disappeared. The graph has lost its record of it. That was the risk in the broader patch, not a security incident I observed. Continuing ingestion would be a poor trade if it made the resulting inventory less trustworthy.

One table-detail request stopped ingestion

The original report described an external table backed by a Delta Lake path in GCS without valid Delta log files. When ingestion called tables().get, BigQuery returned HTTP 400 with the reason invalidQuery. The exception propagated out of ingestion and aborted the sync.

My regression tests inject that status-and-reason combination using the existing test helper:

    error = _make_http_error(400, "invalidQuery")

The tests replace each handler module's gcp_api_execute_with_retry with a function that raises the error. On the pre-fix code, the table-detail handler re-raises it rather than returning None. This reproduces the handler failure without requiring a broken external table or running cloud ingestion.

Recognizing the error was only half the fix

There were two separate checks involved. The shared utility classified the exception, and each handler decided which categories it could tolerate.

Before my change, the classifier's HTTP 400 branch was:

    if status == 400:
        reason = get_error_reason(e)
        if reason.lower() in ("invalid", "badrequest"):
            return "invalid"
        return "unknown"

invalidQuery became unknown. Adding invalid to a handler's tolerated categories would therefore have done nothing by itself. The classifier needed to recognize the reason too.

I checked the other GCP modules named in the report, including storage, compute, and secrets manager. They already handled the invalid category. That comparison supported changing the shared classification, but it did not establish that every BigQuery handler should use the same policy.

The classifier change was small:

-        if reason.lower() in ("invalid", "badrequest"):
+        if reason.lower() in ("invalid", "badrequest", "invalidquery"):

The harder part was deciding what should happen after classification.

The automated review caught what happened after the exception

My initial approach added invalid tolerance across the BigQuery handlers. The automated reviewer @cubic-dev-ai[bot] flagged the table-list path: if listing tables for a dataset failed and ingestion continued, project-scoped cleanup could remove previously ingested table nodes that the failed request had not returned.

I checked the surrounding sync code and agreed with that concern. A handler returning None affected more than its immediate caller. It changed the data that would reach loading and cleanup.

get_bigquery_routines and get_bigquery_connections had the same relevant shape: collection fed cleanup that ran at project scope. I reverted the new tolerance in those functions as well as get_bigquery_tables, then recorded the reasoning in the review thread.

The bot's comment identified one path. Checking the sibling handlers was still my responsibility. Fixing only the function named in the comment would have left the same mistake in the other two.

Failed enrichment can preserve an asset; failed enumeration cannot establish absence

The table-detail path was safe to tolerate for the reported error because the table already existed in the result from tables.list. The merged sync code fetches details and only applies them when the request succeeds:

                detail = get_bigquery_table_detail(client, project_id, dataset_id, tid)
                if detail is not None:
                    table.update(detail)

Returning None here leaves the listed table available for loading. It may lack extra fields such as row and byte counts, but the failed enrichment request does not remove the table from the collected inventory.

Dataset handling had its own guard. sync_bigquery_datasets only loads and cleans up datasets inside the branch where get_bigquery_datasets returns something other than None. If dataset collection fails with this tolerated error, that branch does not run. Dataset detail failure also leaves the listed dataset available.

The final policy was therefore specific:

Handler Behavior for HTTP 400 / invalidQuery
get_bigquery_datasets Return None; skip dataset loading and cleanup
get_bigquery_dataset_detail Return None; retain the listed dataset
get_bigquery_table_detail Return None; retain the listed table
get_bigquery_tables Re-raise
get_bigquery_routines Re-raise
get_bigquery_connections Re-raise

For table details, the final handler change was:

-            ("api_disabled", "billing_disabled", "forbidden", "not_found"),
+            ("api_disabled", "billing_disabled", "forbidden", "invalid", "not_found"),

I did not add the same category to the three list handlers. Retaining those exceptions prevents this fix from letting their invalidQuery failures continue into cleanup with incomplete results.

Regression tests need to preserve the errors we still want

I added invalidQuery to the classifier's parametrized HTTP 400 tests. I also split the handler cases into a group that returns None and a group that must raise.

The raising test uses the same injected error as the tolerant test, but its assertion requires the exception to survive:

    with pytest.raises(HttpError):
        func(client, *args)

That assertion matters because a future edit that broadly adds invalid tolerance could otherwise look reasonable while restoring the cleanup risk. The tests document which entry points can continue and which must stop.

These tests exercise the handler policy; they do not execute Neo4j cleanup. The reason for each policy comes from the sync code, and the tests keep that decision visible when someone changes the handlers later.

The fix shipped in Cartography 0.139.0

@jychp approved the PR before it merged on July 6, 2026. @thenadz also validated the patch against their environment and confirmed that the previously failing cases no longer crashed their scan. That gave the fix a check against the reported cloud behavior alongside my mocked regression tests.

The change shipped in Cartography 0.139.0. Running git tag --contains on merge commit e4c08bd71380d165981f99a2d69982a4846bc134 confirms that this is the earliest containing release tag; the release notes also list the PR.

The lesson I took from this contribution is to trace an error-handling change through every operation that consumes its result. In an inventory system, failed collection does not establish that a resource disappeared. Before returning an empty result or None from an exception handler, check what the caller will infer from it and whether cleanup will still run.

Links

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.