LLM-Assisted Security Testing: Payloads, Code Review, and Detection Notes
Where LLMs fit in a manual test: payload drafting, code review assistance, and detection support, with limits.
LLM-Assisted Security Testing: Payloads, Code Review, and Detection Notes#
Why this topic matters: LLM tools now sit inside many testers' daily workflows, inside the editor, inside the proxy, inside the terminal. Knowing exactly where they help and where they mislead keeps a manual test honest.
Ethics and scope: every note below assumes a written authorization for the system under test. Nothing here covers running real attacks against live third-party targets, and no client data belongs in a public model API. For the human side of web testing this is meant to support, see web hacking basics.
Where an LLM Fits in a Manual Test#
An LLM is a drafting and review aid. It does not hold state, it does not keep your authorization, and it does not observe the target. The tester keeps all three.
What it can do well#
- Draft candidate payloads from a bug class description.
- Explain a piece of application code in plain language.
- Rewrite raw findings into a report paragraph.
- Suggest detection ideas for a vulnerability you already found.
What it cannot do#
- Verify that a payload actually works against a target.
- Know your authorization scope.
- See responses, cookies, or timing unless you paste them in.
- Replace a proxy, a debugger, or a second tester's review.
Payload Drafting from a Bug Class#
The useful pattern is a short prompt that names the context. A payload that works in one sink usually fails in another, so the prompt must name the sink.
Context: JSON API response rendered into an HTML page by a JS framework. Goal: draft candidate strings that might trigger script execution only if the output is inserted without encoding. Format: one string per line, no commentary.
Why the context matters#
A reflected XSS sink, a stored XSS sink, and a DOM sink each need different characters and different escaping assumptions. The same string rarely survives all three. Notes on those differences live in XSS types compared.
Reviewing the output#
Treat every candidate as untested until a request proves it. A model that writes a working-looking string is not a finding; a replayed request that reflects the string into an executing sink is.
Code Review Assistance#
Pasting a small function and asking for a review works best when you give the model the surrounding contract: where the input comes from, what the sink expects, and which framework handles escaping.
Language: PHP 8 Input source: $_GET['id'], used after a intval-like check that was removed in a recent refactor. Sink: concatenated into a SQL string, then mysqli_query. Task: list the exact lines at risk, propose a PDO rewrite, and note any other sink in the same file pattern.
What a good answer looks like#
- It points at specific lines.
- It proposes code you can compile and diff.
- It admits uncertainty when the contract is unclear.
What to distrust#
- Confident claims about a framework's escaping behavior without a version.
- Invented function names that do not exist in the codebase.
- "Best practice" language without a concrete fix.
- Version claims that are not checked against the vendor advisory.
Detection Support#
LLMs are useful for turning a confirmed finding into detection ideas: what a probe request looks like in logs, what a WAF rule would match, and which response header suggests the sink is unescaped.
| Finding | Signal in logs | Signal in response |
|---|---|---|
| Reflected XSS probe | Unusual characters in query params | Payload echoed unescaped in HTML |
| SQL error-based probe | Repeated 500s on one parameter | Database error text in body |
| Second-order payload | Stored value containing markup | Execution when admin views list |
| Blind XSS callback | Inbound request from a browser UA | None; observed on callback server |
Keep detection separate from exploitation#
Write detection notes from the confirmed request and response. Do not let the model invent log fields the target does not emit. Compare your notes against the response fields a proxy shows; Burp Suite basics covers that loop.
Privacy and Data Handling#
Client code, session cookies, and internal hostnames should not leave your machine without an agreement that covers the model provider. Practical rules:
- Paste the minimum snippet needed for the question.
- Replace real hostnames, real usernames, and real tokens with placeholders.
- Prefer local or on-premises models for sensitive code.
- Check whether your engagement contract names approved tools.
Model Failure Modes to Expect#
- Hallucinated CVEs and invented library versions. Always verify against the vendor advisory.
- Confident but stale advice about framework defaults. Pin the version in your prompt.
- Echoing your prompt back as if it were a finding. Ask for the exact lines and sink.
- Overclaiming exploitability from dead code paths. Trace the call chain yourself.
Tooling Notes#
Several editors and proxies now ship LLM integrations. The useful ones in a test are the ones that read your actual traffic: proxy history, repeater tabs, and local notes. The risky ones are the ones that upload full projects silently. Check the setting before the first run.
# Example: piping a finding note into a local model CLI $ cat finding-notes.md | llm "Summarize impact and draft the request appendix"
No vendor names are required for the note to be useful. Choose the model that matches your data policy, not the one with the loudest marketing.
How This Fits With Other Skills#
- Payload hygiene: WAF bypass and Unicode shows why prompts must name the sink.
- Injection classes: SQL injection notes lists the sinks a model will blur together.
- Secondary contexts: attacking secondary contexts explains sinks an LLM cannot see at all.
- Test planning: penetration testing roadmap orders these skills.
Limitations#
- This note reflects practitioner experience and vendor documentation, not a controlled study.
- No statistics are cited here because tool reliability changes month to month.
- Model behavior varies across versions, prompts, and context lengths.
- Nothing above is legal advice. Authorization scope always comes from your contract.
Suggested Workflow#
- Find the candidate sink manually, with the proxy.
- Draft candidates with the model, naming context and encoding assumptions.
- Replay candidates with a request; keep only the ones that execute.
- Write the finding from the request and response, not from the model's summary.
- Add a detection note while the traffic is still in the proxy history.
Closing Notes#
LLMs speed up drafting and review. They do not remove the need to replay, verify, and document. Treat the model as a fast pair programmer with no memory, no scope, and no access to the target. The findings you ship should still stand on their own request and response evidence. For tool reviews that stay grounded in observed behavior, see ZAP 2.16 review.
A Worked Review Session#
A short example shows the shape of a useful session. The snippets are illustrative, not output from a real engagement.
Step 1: paste a minimal sink#
$id = $_GET['id']; $row = mysqli_query($conn, "SELECT * FROM items WHERE id=$id"); echo $row['name'];
Step 2: ask for a classification, not a verdict#
Classify the flaw on line 3. Name the CWE class, the sink, the source, and list what a parameterized rewrite looks like. Do not speculate about reachability.
Step 3: verify with a request#
Only a replayed request with a visible error or a time difference turns the classification into a finding. A model's confidence is not evidence.
Step 4: write it up from the traffic#
Keep the proven request, the trimmed response, and the exact parameter in the report. The model's summary can be a draft, but the request log is the record.
Reviewer Checklist for LLM-Assisted Notes#
Run this list before a finding leaves your machine:
| Item | Pass condition |
|---|---|
| Sink named | The exact parameter and function appear |
| Request logged | At least one working request is attached |
| Payload tested | Every quoted string was replayed once |
| Scope stated | The note says which host it applies to |
| Privacy kept | No real client tokens remain in the note |
| Detection added | One log signal and one response signal listed |
| Limitations recorded | What was not tested is written down |
Common Mistakes in Practice#
- Letting the model name the bug class from a stack trace alone.
- Shipping the model's impact paragraph when the request never fired.
- Pasting a full client project into a public API "for context".
- Trusting a version number in a suggestion without checking the advisory.
- Forgetting that a suggested payload may simply not encode for the sink you have.
When Not to Use an LLM at All#
- For anything covered by the engagement's data rules.
- When the sink is already clear and a one-line grep answer would do.
- When a second human can review faster than you can verify the output.
- When the client contract names approved tooling explicitly.
Prompt Templates Worth Reusing#
Short, specific prompts get the model into the right shape of answer. Long vague ones get a wall of text that still misses the sink.
Template: classification request#
Classify the flaw at the marked line. Give: bug class, source, sink, one-line impact hypothesis. Do not suggest exploitation steps. <paste 15 lines of code around the mark>
Template: payload drafting#
Context: HTML comment context inside a stored profile field. Constraint: output is filtered for angle brackets. Ask: candidate strings that survive bracket filtering and still break out of the comment.
Template: report rewrite#
Rewrite the following finding as three short paragraphs: impact, evidence, fix. Keep only facts present in the notes. Do not add CVSS, do not add claims.
Template: detection draft#
Given this request and response pair, list what a SIEM rule could match on, and which match would have the fewest false positives on a normal login flow.
Reading Order for the Rest of This Series#
- First principles of web classes: web hacking 101.
- Injection sinks in detail: SQL injection notes.
- Markup injection: XSS types compared.
- Traffic tooling that grounds the notes: Wireshark guide.
- Hidden sinks an LLM cannot see: attacking secondary contexts.
- Unicode and normalization traps: bypassing WAFs with Unicode.
- Client-side object traps: prototype pollution.
- WordPress sinks worth checking: WordPress attack surface.
- Cloud misconfiguration checks: AWS testing notes.
- Mobile transport checks: mobile app testing notes.
- Contract flaws in Solidity: smart contract auditing.
- Low-privilege pivots in cloud IAM: AWS testing notes.
- Human-side attack patterns: social engineering patterns.
- Zero-interaction sinks: zero-interaction XSS.
- Exploit stages at a high level: exploit development stages.
- Metasploit task map: Metasploit notes.
- Core risk categories: OWASP Top 10 notes.
- Toolkit overview: bug bounty toolkit.
- Roadmap context: penetration testing reading order.
- Dead references as risk: dead link detection.
- The proxy workflow underneath all of this: Burp Suite basics.
Hygiene Checklist Before Pasting#
| Field | Action |
|---|---|
| Hostnames | Replace with target.local |
| Tokens | Replace with REDACTED |
| Usernames | Replace with user-a |
| Internal IPs | Replace with 192.0.2.1 |
| Client names | Remove entirely |
| Dates | Keep, they help later review |
Appendix: What "Verified" Means Here#
A claim in this note is either tied to a request you can replay, tied to a vendor document you can name, or marked as experience. Anything else is a draft, not a finding. That three-way split is the whole discipline, and it is the part no model can do for you. A working replay, a named source, or an honest label. Pick one before the note leaves your machine.
Two replay rules carry over from the rest of testing. First, a string that looks dangerous but never executes is a curiosity, not a bug. Second, a bug you cannot replay twice is a hint, not a finding. Both rules apply doubly when the string came from a model, because model output looks authoritative by default. Keep the receipts, and the notes stay useful.
Keeping the Notes Durable#
Model output ages badly. Store the request, the response, and the exact prompt you used in the same folder as the finding. A year later the request still proves the bug; the model's prose may not even parse. That habit also makes dead link and reference checks easier, because the evidence lives with the claim. Version every note with the date and the model or tool version used. When a model ships a new version, re-run one archived prompt and diff the output. When a tool updates its UI, check that the setting you relied on still exists. Small habits, kept over months, are what separate durable notes from a pile of screenshots.
What do you think?
React to show your appreciation