Shadow APIs and Asset Inventory: Mapping the Attack Surface

You cannot protect an endpoint you do not know: an inventory record format and coverage metric built from code, traffic, and DNS sources.

21 min read
ibrahimsql
4,081 words

Shadow APIs and Asset Inventory: Mapping the Attack Surface#

An endpoint that is not on the test list does not count as tested. Most teams judge API security by looking at a few known routes: login, user profile, payment, admin panel. The production attack surface is wider than that. Old versions, routes opened for internal tools, hidden parameters used by the mobile release, and routes that never reached the documentation live on the same network.

This post describes managing the attack surface with records instead of guesses. The goal is not a single list, but showing how the list is produced, how it stays current, and how it connects to test coverage. An inventory is not a spreadsheet; it is a dataset fed by code and traffic.

The term shadow API is used here in a narrow sense: every HTTP interface that receives requests but has no inventory record. That includes forgotten version paths, internal debug endpoints, mobile-only endpoints, and third-party callbacks. Without a record there is no owner, and without an owner there is no test. The order of the post is as follows: why inventory comes before testing, the six sources that feed it, the record format, the diff automation, version discipline, the coverage metric, and finally the limits.

1. Why inventory first: unknown endpoints stay unprotected#

Authorization testing runs against a known endpoint list. An incomplete list produces an incomplete result, silently. The report says "no critical findings" because there are places the report never looked. That condition is the shadow API problem, and it appears in three patterns.

1.1 Forgotten version paths#

Version two ships, version one waits to be shut down, and the wait lasts for years. The old version usually carries the old authorization logic with it: the new tenant filter is written into the version two layer, while the checker on the version one side keeps working with old logic. For an attacker the version number is a detail; every path that accepts requests is an entrance.

This pattern grows during gradual migrations. Old mobile releases keep calling version one, and the team cannot close the old path in the name of backward compatibility. When the path that cannot be closed has no record, nobody asks whether it sits in the test matrix either.

1.2 Internal debug and support routes#

Endpoints opened to look things up fast during an incident are typical: user search, order detail, bulk export. They start with an IP restriction or a temporary key, then the restriction loosens, the key ends up in code, the route is forgotten. These routes never reach the documentation because they do not count as product features.

The trouble is that these routes show more data than ordinary ones. Debug output carries many fields: related records, internal notes, raw states. Missing access checks therefore expose more data. The object-level checks from the API access control test plan are exactly the kind of checks that break on such paths.

1.3 Mobile-only endpoints#

The web team and the mobile team seem to share one API, but in practice the mobile application uses extra parameters and extra paths. Device registration, push notification confirmation, in-app purchase verification never appear in the web flow. A team that builds its inventory from web traffic misses them. Mobile endpoints also carry identity differently: long-lived refresh tokens, device-bound keys, receipt verification through the store. A test plan written around web sessions leaves mobile endpoints out.

1.4 A concrete scenario: the undocumented export endpoint#

Consider this flow: a shop admin panel has an order list page, the page calls /v2/orders, and testing focuses there. The same panel once needed a bulk export feature, so a developer added /v1/orders/export?format=csv. The feature was removed from the interface, but the route stayed in code.

Months later the path still receives requests. An old mobile release calls it for background sync, and a few saved scripts keep calling it too. The new tenant filter was written into the version two layer, while the checker on the version one side runs with old logic. The result: the tested path looks safe, and the untested path serves the same data through a different door.

The lesson of the scenario is direct: security decisions are made per endpoint, not per screen. A feature removed from the screen does not count as removed from the API. The record format and diff automation sections exist to close that gap.

Shadow patternTypical originWhy testing skips itFirst sign
Forgotten /v1Gradual version migrationTests target the current versionRequests to the old path in logs
Debug routeIncident toolingNever enters documentationPath missing from docs
Mobile-only endpointApplication releasesScanning with web trafficExtra paths in the bundle
Support panel endpointInternal toolsNot counted as productAPI on a subdomain
Webhook receiverThird-party integrationInbound direction forgottenUnsigned requests accepted
Batch job endpointScheduled tasksInteractive tests never see itTraffic in nightly logs

2. Inventory sources#

No single source produces a full inventory. What the code says differs from what the traffic says, and what DNS says is different again. The six sources below complete each other: each catches something different and misses something different.

2.1 Route-file scanning#

Source code mirrors intent. Route definitions tell which paths may exist. The scripted approach is simple: walk the route files, collect path patterns from decorators and registration calls, pair them with the HTTP method, and write the output as a list.

# First pass: roughly collect route definitions rg -n --no-heading \ -e "@(app|router)\.(get|post|put|patch|delete)" \ -e "(app|router)\.(get|post|put|patch|delete)\(" \ -e "path\s*=\s*[\"']/api" \ src/ | sort -u > /tmp/route-hits.txt wc -l /tmp/route-hits.txt head -n 20 /tmp/route-hits.txt

A small Python script gives a tidier search. It parses per file extension and reports path and method in separate columns. The aim is a fast candidate list, not perfect parsing.

# route_scan.py: build a candidate endpoint list from route files import pathlib import re PATTERNS = [ re.compile(r"""@(?:app|router)\.(get|post|put|patch|delete)\(\s*["']([^"']+)["']"""), re.compile(r"""(?:app|router)\.(get|post|put|patch|delete)\(\s*["']([^"']+)["']"""), re.compile(r"""path\s*=\s*["']((?:/api|/v\d|/internal|/debug)[^"']*)["']"""), ] ROOTS = ["src/routes", "src/api", "app/api", "server/routes"] def scan(root: pathlib.Path): rows = [] for file in root.rglob("*.py"): text = file.read_text(encoding="utf-8", errors="ignore") for rx in PATTERNS: for m in rx.finditer(text): if len(m.groups()) == 2: rows.append((m.group(1).upper(), m.group(2), str(file))) else: rows.append(("?", m.group(1), str(file))) for file in root.rglob("*.ts"): if "node_modules" in str(file): continue text = file.read_text(encoding="utf-8", errors="ignore") for rx in PATTERNS: for m in rx.finditer(text): if len(m.groups()) == 2: rows.append((m.group(1).upper(), m.group(2), str(file))) else: rows.append(("?", m.group(1), str(file))) return sorted(set(rows)) if __name__ == "__main__": found = [] for r in ROOTS: p = pathlib.Path(r) if p.exists(): found += scan(p) for method, path, file in found: print(f"{method:6} {path:45} {file}") print(f"\nTotal candidates: {len(found)}")

What this source reliably catches is summarized in the points below:

  • Every path defined in code, documented or not

  • Method data: GET and POST variants of one path become separate rows

  • File location: hints at who the owner might be What this source still misses is listed in the points below:

  • Paths added at the gateway layer never appear in code

  • Dynamic paths produced by plugins and middleware

  • Paths present in code but disabled in deployment (dead-code noise)

  • Unusual registration styles of different frameworks

2.2 OpenAPI document against reality#

The OpenAPI document is how the team sees the world. Real traffic is how the world actually is. The gap between the two is the most productive inventory input. The comparison runs both ways: in the document but absent from traffic (dead or wrong docs), in traffic but absent from the document (shadow candidates).

# Quickly count paths in the OpenAPI document python3 -c " import json spec = json.load(open('openapi.json')) paths = spec.get('paths', {}) print(f'Documented paths: {len(paths)}') for p in sorted(paths)[:15]: print(' ', p) "

What the document side reliably catches is summarized in the points below:

  • Schema data: parameters, bodies, response shapes

  • Identity expectations: the security requirement of each operation

  • Version data: which path belongs to which version What the document side still misses is listed in the points below:

  • Everything, when the document is stale: documents tend to rot

  • Internal paths kept out of the document

  • Variants hidden behind wildcards in the document

  • Mobile endpoints that never entered the document

2.3 Proxy history review#

Proxy history shows what actually gets called. Traffic remembers what the code forgot. A proxy log from a test or production environment reveals paths missing from the document. The aim is not reading single requests, but counting and ranking path patterns.

# Extract path frequency from a proxy export python3 -c " from collections import Counter import re c = Counter() with open('proxy-history.txt') as f: for line in f: m = re.search(r'\s(GET|POST|PUT|PATCH|DELETE)\s(\S+)', line) if m: method, url = m.groups() path = re.sub(r'^https?://[^/]+', '', url).split('?')[0] c[(method, path)] += 1 for (method, path), n in c.most_common(30): print(f'{n:6} {method:6} {path}') "

What the proxy source reliably catches is summarized in the points below:

  • Paths that truly receive requests, free of dead code

  • Parameter and body samples

  • Whether identity headers are actually carried

  • Hidden paths that return errors (404 and 403 rows are signals too) What this source still misses is listed in the points below:

  • Rare paths never called during the recording window

  • Endpoints run by scheduled jobs at night

  • Paths triggered only under specific tenants

  • Endpoints served from another region

2.4 Endpoint extraction from JS bundles#

The frontend bundle is the API client's confession. Compiled JS keeps path strings inside. String scanning collects /api/... patterns, then intersects them with traffic and documentation. A path found in the bundle and nowhere else deserves a look.

# Search path strings in the compiled bundle rg -o --no-filename \ -e '"/api/[^"]+"' \ -e "'/api/[^']+'" \ -e '"/v[0-9]+/[^"]+"' \ -e "'/v[0-9]+/[^']+'" \ dist/ web/.next/static/ 2>/dev/null \ | tr -d "\"'" | sort -u > /tmp/bundle-paths.txt wc -l /tmp/bundle-paths.txt head -n 20 /tmp/bundle-paths.txt

What the bundle source reliably catches is summarized in the points below:

  • Paths the frontend truly calls

  • Client endpoints forgotten by the docs

  • Paths called conditionally behind feature flags

  • Extra paths left in debug builds What this source still misses is listed in the points below:

  • Server-side calls (scheduled jobs, queue consumers)

  • Paths used by the mobile application

  • Obfuscated or piecewise-concatenated strings

  • Third-party callback endpoints

2.5 DNS and subdomain enumeration#

An API sometimes lives not on the main domain but on a forgotten subdomain: records such as admin-old, api-test, or support-panel live for years. DNS enumeration finds these doors. Certificate records and passive DNS data list the subdomains; each candidate is then probed for what answers on it.

# Collect subdomains from certificate records curl -s "https://crt.sh/?q=%25.example-company.com&output=json" \ | python3 -c " import json, sys try: data = json.load(sys.stdin) except Exception: data = [] names = set() for row in data: for part in str(row.get('name_value', '')).split(): names.add(part.strip().lower()) for n in sorted(names)[:40]: print(n) print(f'Total candidate subdomains: {len(names)}') "

What the DNS source reliably catches is summarized in the points below:

  • Forgotten test and staging environments

  • Old admin panels

  • Internal tools left exposed

  • Regional or tenant-specific endpoints What this source still misses is listed in the points below:

  • Path-level detail (it yields hosts only)

  • Identity and data-class data

  • Internal endpoints never published to DNS

  • Short-lived campaign subdomains once closed

2.6 Mobile API traffic#

The mobile application is a separate client and is watched separately. Application traffic is routed through a proxy from a test device, with certificate pinning relaxed in the test build. The collected paths are intersected with the web inventory: every path outside the intersection is a mobile-only candidate.

What this source reliably catches is summarized in the points below:

  • Endpoints only the application calls

  • Version parameters and device headers

  • Mobile flows such as store receipt verification

  • Paths old application releases still call What this source still misses is listed in the points below:

  • Paths of old releases nobody installs anymore

  • Endpoints rarely triggered by background sync

  • Server-to-server calls

  • Dead paths embedded in the device build but no longer called The summary table collects the sources. It shows which source closes which gap.

SourceMain inputStrongest sideBlind spot
Route scanningSource codeFinds undocumented pathsDead-code noise
Document comparisonOpenAPISchema and security dataMisleads when stale
Proxy historyReal trafficPaths actually usedMisses rare paths
JS bundleCompiled frontendClient-side truthBlind to server jobs
DNS enumerationCertificates and DNSForgotten subdomainsNo path detail
Mobile trafficDevice proxy logMobile-only endpointsMisses old releases

3. Inventory record format#

A list rots without a record format. The same fields are filled for every endpoint, and an empty field counts as data too: an endpoint with an unknown owner is an unowned endpoint. The schema below gives the smallest sufficient set.

FieldRequiredMeaningExample
methodYesHTTP methodGET
pathYesPath as a pattern/v1/orders/{id}/export
auth_typeYesHow identity is carriedbearer, cookie, api-key, none
tenant_scopeYesWhether tenant separation existsper-tenant, global, unknown
data_classYesSensitivity of carried datapublic, internal, personal, payment
ownerYesResponsible team or personpayments-team
statusYesLifecycle stateactive, deprecated, unknown
versionNoVersion labelv1
sourceNoWhich source found itproxy, route-scan, bundle
last_seenNoLast observation date2026-09-28
pii_fieldsNoPersonal data fields["email", "iban"]
notesNoContext noteused by the support panel
The meaning of the fields, briefly. Every record with auth_type set to none gets a separate review: the endpoints that must be public are known (health checks and similar), any other none record is a candidate finding. Records with tenant_scope set to unknown take priority in testing, because tenant separation stands unverified.

The data_class field sets the test order. Endpoints carrying payment and personal data go first, endpoints serving public content go later. That ordering points the same way as impact-first thinking: close the places with the largest harm first.

The status field takes three values; two are not enough. active marks a supported path, deprecated a path in shutdown, unknown a path nobody has confirmed yet. Every newly found record enters as unknown and stays there until the owning team confirms it.

An example record for this schema looks like this:

{ "method": "GET", "path": "/v1/orders/{id}/export", "auth_type": "bearer", "tenant_scope": "unknown", "data_class": "personal", "owner": "unknown", "status": "unknown", "version": "v1", "source": "proxy", "last_seen": "2026-09-28", "pii_fields": ["email", "phone", "address"], "notes": "Seen in proxy log, missing from docs. Owning team to confirm." }

A second example shows a healthy record. All fields are filled and the state is clear.

{ "method": "GET", "path": "/v2/orders/{id}", "auth_type": "bearer", "tenant_scope": "per-tenant", "data_class": "personal", "owner": "orders-team", "status": "active", "version": "v2", "source": "openapi", "last_seen": "2026-10-01", "pii_fields": ["email", "phone"], "notes": "Tenant filter applied server side, present in the test matrix." }

Records live in a single JSON Lines file. Each row is one endpoint and enters version control. The arrangement pays off plainly: the diff automation reads the file, the metric is computed from it, and review requests open against it.

# Basic health check of the inventory file python3 -c " import json rows = [json.loads(l) for l in open('api-inventory.jsonl') if l.strip()] print(f'Total records: {len(rows)}') from collections import Counter print('Status split:', dict(Counter(r['status'] for r in rows))) print('Unowned records:', sum(1 for r in rows if r.get('owner') == 'unknown')) print('Unknown scope:', sum(1 for r in rows if r.get('tenant_scope') == 'unknown')) print('No-auth records:', sum(1 for r in rows if r.get('auth_type') == 'none')) "

4. Spec-vs-traffic diff automation#

The document stays fresh only when docs and traffic meet regularly. The approach has three steps: read the document, read the traffic, report what falls outside the intersection. Comparison happens at path-pattern level, not raw-URL level. /orders/123 and /orders/456 are one pattern: /orders/{id}.

Paths are normalized first. Numbers, UUID shapes, and date forms turn into placeholders. Without this step every request looks like a separate path and the report becomes noise. Normalization rules stay small and live in a file.

Document paths and observed paths are then intersected. Every pattern present in traffic but missing from the document is a shadow candidate. Every pattern present in the document but missing from traffic is dead or wrong documentation. Both directions get reported, because both carry maintenance signal. The pseudocode below shows the flow. It reads like TypeScript, aims at being a production script skeleton, and is not meant to run as shown.

// diff.ts: compare OpenAPI paths with proxy observations type ObservedHit = { method: string; rawPath: string; count: number }; type Finding = { kind: "undocumented" | "unstaged-doc" | "method-mismatch"; method: string; path: string; count: number; }; const ID_LIKE = [ /\/\d+(?=\/|$)/g, /\/[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}(?=\/|$)/gi, /\/\d{4}-\d{2}-\d{2}(?=\/|$)/g, ]; function normalize(rawPath: string): string { const clean = rawPath.split("?")[0].replace(/\/+$/, "") || "/"; let out = clean; for (const rx of ID_LIKE) { out = out.replace(rx, "/{id}"); } return out; } function openApiKeys(spec: any): Set<string> { const keys = new Set<string>(); const paths = spec.paths ?? {}; for (const [path, ops] of Object.entries<any>(paths)) { const normalized = normalizeSpecPath(path); for (const method of Object.keys(ops)) { if (["get", "post", "put", "patch", "delete"].includes(method)) { keys.add(method.toUpperCase() + " " + normalized); } } } return keys; } function normalizeSpecPath(specPath: string): string { return specPath.replace(/\{[^}]+\}/g, "{id}"); } function diff( spec: any, hits: ObservedHit[] ): Finding[] { const documented = openApiKeys(spec); const seen = new Map<string, number>(); for (const h of hits) { const key = h.method.toUpperCase() + " " + normalize(h.rawPath); seen.set(key, (seen.get(key) ?? 0) + h.count); } const findings: Finding[] = []; for (const [key, count] of seen) { const [method, ...rest] = key.split(" "); const path = rest.join(" "); if (!documented.has(key)) { findings.push({ kind: "undocumented", method, path, count }); } } for (const key of documented) { if (!seen.has(key)) { const [method, ...rest] = key.split(" "); findings.push({ kind: "unstaged-doc", method, path: rest.join(" "), count: 0, }); } } return findings.sort((a, b) => b.count - a.count); }

The script output stays readable for humans. Each row asks for one decision: add to inventory, fix the document, or close the path. Undecided rows enter the inventory as unknown and get an owner assigned.

# diff-report.txt (sample output) UNDOCUMENTED GET /v1/orders/{id}/export 412 requests UNDOCUMENTED POST /internal/support/lookup 88 requests UNDOCUMENTED GET /api/mobile/receipt/verify 61 requests UNSTAGED-DOC DELETE /v2/orders/{id} 0 requests UNSTAGED-DOC PATCH /v2/users/{id}/role 0 requests

The run rhythm matters. It runs once per release plus as a weekly scheduled job. A threshold is set: an undocumented path above, say, 10 requests opens a review ticket automatically. Below the threshold stays watched quietly, above it reaches a human.

5. Version and deprecation handling#

The forgotten-/v1 problem is a process problem, not a technical one. The shutdown decision is never made, only delayed, and delay turns permanent. The fix defines shutdown as a four-stage job instead of a single event: announce, measure, block, remove. Each stage leaves its mark on the inventory record.

The announce stage flags the path as deprecated. A sunset header goes on responses, and client owners get a date. The header is a small job with a large effect: logs reveal which clients still use the old path.

# Response headers during shutdown (example) Deprecation: true Sunset: Wed, 01 Apr 2026 00:00:00 GMT Link: </v2/orders/{id}>; rel="successor-version"

The measure stage watches usage. Which client release, which tenant, how many requests per day: these numbers show whether the shutdown date is realistic. When usage refuses to reach zero the date is not postponed; blocking is planned instead. Postponement without measurement is not allowed.

The block stage closes the path by default. Clients that need an exception pass through an allow list temporarily. Every exception lands in the inventory note with an end date. Exceptions without an end date are not accepted.

The remove stage deletes the code, deletes the test, and flips the record to removed. The record is not deleted from the file; its state changes: the knowledge that such a path once existed is kept. During review the question "what happened to this path" already has an answer.

The checklist for the shutdown process is as follows:

  • Marked deprecated in the inventory, date written
  • Response headers added (Deprecation, Sunset)
  • Client owners notified of the date
  • Daily usage counting started
  • Migration plan made with above-threshold users
  • Block date written into the inventory
  • Allow-list exceptions recorded with dates
  • Code and tests deleted, record set to removed

6. Coverage metric: share inside the test matrix#

The inventory earns its keep when tied to testing. The metric is one sentence: what share of inventoried paths has entered the authorization test matrix. The denominator is the inventory, the numerator the tested records.

coverage = inventory_records_in_test_matrix / total_active_inventory_records

Example: with 180 active records where 135 appear as rows in the matrix, coverage is 0.75. The target number is a management call; this post suggests no threshold. What matters is recomputing the number on every release.

The metric splits by data class. Endpoints carrying payment data and endpoints serving public content sit on separate rows of the same table. One average hides a critical gap: the overall rate can look high while payment endpoints stand empty.

SliceActive recordsIn matrixCoverage
Payment data22200.91
Personal data64480.75
Internal data51300.59
Public43370.86
Total1801350.75
Tracking happens per release. The metric for each release tag is stored as one row, and a drop gets an explanation in the release notes. A drop has two usual causes: new paths were added, or old paths fell out of the matrix. Both read from the inventory; no guessing needed.
# Keep per-release coverage history in a simple file cat coverage-history.csv # version,active,tested,coverage # v2.14.0,172,131,0.76 # v2.15.0,180,135,0.75 # v2.16.0,184,150,0.82

This metric reads together with the matrix from BOLA test matrix automation. The matrix gives the depth of testing, the inventory the width. Combined they answer "what did we test, and how much of it" with two numbers: coverage and the passing-test rate.

7. Limits and Short result#

Inventories tend to rot. Sources change, scripts stop running, DNS entries age, mobile release traffic goes uncollected. The cure is not a one-time cleanup but scheduled work: the diff script runs weekly, coverage is computed per release. A skipped step turns visible, and the visible gets fixed.

Client-side endpoints stay a partial blind spot. Obfuscated bundles, piecewise-joined strings, and paths left in local builds may never reach the full list. The gap is accepted and written down: the inventory keeps a "known unknowns" section too.

Third-party callbacks need separate care. The inbound endpoint hides in the shadow of the outbound integration. A receiver without signature verification is handled like a no-auth record. When the provider side changes, the path starts behaving differently without notice, so receiver tests are read together with provider release notes.

Short result: the unprotected endpoint is usually the unknown endpoint. Six sources merge into one record format, diff automation keeps the list current, and the coverage metric turns "how much of it did we see" into a number. An inventory is not a finished document; it is a dataset renewed each release.

---
Share this post:

What do you think?

React to show your appreciation

Related Posts