Building Self-Reviewing Code Generation Loops for Web Extraction

03 October 2026 | 22 min read

Self-reviewing code generation for web extraction means letting an LLM write real extraction code, running that code against real page HTML, checking the output mechanically, and revising the code when verification fails.

In this article, we'll build that loop ourselves instead of treating generated code as correct by default. We'll start with a simple web scraping test against books.toscrape.com, then make the setup harder with holdout pages, JavaScript-rendered content, and infinite scroll.

Along the way, we'll look at what actually failed in our experiments, how we separated parser problems from fetch problems, and where ScrapingBee fits as the fetch layer without pretending that it generates or repairs the extraction code for us.

Building Self-Reviewing Code Generation Loops for Web Extraction

TL;DR

  • The loop has four stages: Generate, Execute, Verify, Refine (GEVR).
  • It is not a self-healing scraper: the model writes and revises real source code instead of simply relocating broken selectors.
  • ScrapingBee's ai_query and ai_extract_rules are different: they return extracted data directly, with no reusable code artifact or review loop.
  • In my tests, the Books extractor passed immediately, JavaScript rendering changed 0 product cards into 16, and infinite scroll required a fetch refinement from 6/48 items to the full 48/48.
  • Cap Refine at roughly two or three attempts, then fail loudly and hand the evidence off for review.
  • ScrapingBee's MCP server can provide the fetch layer for an agent running this pattern; it does not generate, execute, or repair the extraction code itself.

How self-reviewing code generation for web extraction works

A self-reviewing code generation loop is a simple idea: don't trust generated code just because it looks right — run it and check what it actually does.

The pattern can be used for many kinds of LLM-generated code. In this article, I'll use web extraction as the concrete example, because it gives us something easy to test: real page HTML goes in, structured data comes out, and we can check whether that data is actually correct.

Say you ask an LLM to write a Python extractor for a product page. It gives you a neat-looking script, the selectors seem reasonable, and nothing obviously looks broken. The tempting move is to run with it.

A self-reviewing loop adds another layer around that generated code. The script runs against real HTML, its output goes through a separate verification step, and a failed result becomes feedback for the next attempt. In the next section, we'll call these stages Generate, Execute, Verify, and Refine (GEVR).

What problem does this solve?

Generated extraction code has an awkward failure mode: it can run successfully and still return incomplete or subtly wrong data. A broken selector does not always throw an exception; it might quietly return 12 products when the page contains 20. A price parser can grab the wrong number and still produce something that looks completely believable.

That is the main reason to add verification. Instead of treating "the program ran" as success, you check whether the result actually makes sense.

For recurring targets, there is another benefit. Once an extractor passes verification, you can keep the resulting script and run it again without asking an LLM to reinterpret every page. If the site changes later and those checks start failing, that is when the generation loop can come back into play.

Verification can also help you fix the right problem. As our experiments later in the article show, a bad result does not necessarily mean the generated parser is wrong. The page may have been fetched incompletely, or the verification rule itself may be mistaken.

Keep enough evidence to understand a failure

A useful loop leaves a paper trail behind. For each attempt, you usually want to keep:

  • the generated source code;
  • stdout, stderr, and the exit code;
  • the verification result;
  • and, after a revision, a diff showing what changed.

That may sound like a lot for a small scraper, but it becomes useful as soon as a failure is subtle. Suppose the model fixes one broken selector but also rewrites code that was already working. A diff makes that immediately visible: it tells you whether the model made a small repair or decided to redecorate the whole house.

The same records are useful when the loop eventually gives up. Instead of just getting "generation failed," you have the code, execution result, verification errors, and revision history needed to understand why.

New to Python web scraping? See ScrapingBee's Python web scraping tutorial for the basics of fetching pages, parsing HTML, and extracting data before adding a review loop.

What this pattern is not

There are a few nearby ideas that are easy to mix up with this one.

A self-healing scraper usually tries to recover when a selector breaks by finding the same element again from structural or visual clues. That can be useful, but it solves a different problem. In our loop, the artifact being revised is actual source code: the model writes it, the program runs, the result is checked, and the source may be changed after a failure.

This is also different from ScrapingBee's ai_query and ai_extract_rules. Those features let you describe the data you want and get extracted data back from a request. They do not create a reusable Python program and put it through an external execution-and-review loop.

ScrapingBee plays a different role in the experiments later in this article: it provides the page HTML the generated extractor works with, including rendered HTML when JavaScript is involved. The generation, execution, verification, and revision logic belongs to our application.

And when people say the AI is "rewriting its own code," that needs a small asterisk. It is revising source code that it generated earlier. It is not changing its model weights or retraining itself after every failed run. So the whole pattern can still be reduced to one sentence: generate some code, run it for real, check the result, and only then decide whether it is good enough.

The Generate, Execute, Verify, Refine loop: naming it, and the research it draws from

For this article, let's use Generate, Execute, Verify, Refine (GEVR) as a convenient name for the loop.

GEVR is not an established industry or research term but the underlying pattern is familiar: a model produces an output, observes what happened when that output is used, receives feedback, and tries again.

How the four stages work

Each stage has a fairly narrow job:

  • Generate writes a runnable extraction program from the page and extraction requirements.
  • Execute runs that program against real HTML and records what happened.
  • Verify checks the result against rules we define ourselves.
  • Refine gives the failed code and useful feedback back to the model so it can produce another version.

Verify is the important bit. Asking the same model, "Does this code look correct?" is not much of a check. If it made a bad assumption while generating the extractor, it may make the same assumption while reviewing it.

Instead, verification should use evidence from outside that reasoning process: an exit code, stderr, a schema failure, an unexpected record count, a mismatched value, or another deterministic check.

Where the idea comes from

This general approach has a few close relatives in LLM research:

  • Reflexion, introduced by Shinn et al. and published at NeurIPS 2023, has an agent use feedback from earlier attempts as context for later ones. The model itself is not retrained between attempts. Its weights stay the same; the feedback changes what information it has for the next try. That is close to what happens in Refine: the model sees what went wrong and produces another version using that information.
  • Teaching Large Language Models to Self-Debug by Chen et al. gets even closer to generated code. Their Self-Debugging approach uses program behavior and execution feedback when diagnosing and revising an answer. The model is not limited to staring at its own source and guessing whether it works.
  • A more recent example is the 2026 paper ASG-LLM, which applies a similar idea to generated scraping modules. Its pipeline generates code from browser traffic, executes it in a controlled environment, and uses structured failure information when repairs are needed.

That work deals with more security-sensitive and session-bound workflows than the experiments in this article. Here, we'll keep the scope to publicly accessible pages.

Running generated code safely

There is one practical issue that research diagrams tend to make look deceptively simple: generated code still has to run somewhere.

Do not give it unrestricted access to the machine running your application. A production setup might inspect the generated source before execution, restrict allowed imports and operations, remove secrets from the environment, impose time and output limits, and run the process inside an isolated container or similar restricted environment.

Static checks can catch obvious problems, but they are not a security boundary by themselves. They can also produce false positives, so the rules around execution need testing just like the generated extractor does.

Give Refine useful evidence

When execution or verification fails, the next prompt should contain information that helps explain the failure. That can include the previous source code, stderr, relevant execution output, failed verification checks, and a diff from the previous revision.

There is little benefit in feeding an ever-growing transcript of every previous attempt back into the model. Keep the feedback focused on what happened, and put a hard limit on how many repair rounds the loop can perform.

Reuse code that already passed

Once an extractor passes verification, there is usually no reason to generate it again for every request. Keep the verified version and reuse it while the target structure remains compatible. If verification begins failing later, the loop can be invoked again.

This is also why GEVR is different from AI-era scraping tools like Crawl4AI. It is not a library or product. It is an orchestration pattern around code generation, real execution, verification, and controlled revision.

First-hand test: building and running the loop

I wanted to test the pattern with real generated code rather than inventing a convenient failure for the article. I started with Books to Scrape, a public site designed for scraping practice. It is mostly static HTML, so it gave me a simple baseline before moving on to JavaScript and infinite scroll.

For these tests, I used GPT-5.6 through OpenAI's Responses API to generate Python extractors with BeautifulSoup.

How the test runner works

The experiment is a small Python project rather than a full scraping framework:

  • runner.py handles fetching, code generation, execution, retries, and saved artifacts;
  • prompts.py contains the generation and refinement prompts;
  • verifier.py checks the extracted output independently;
  • contract.json describes the fields the generated extractor must return;
  • pyproject.toml and uv.lock define the environment.

The important separation is between fetching, generated code, and verification:

  • ScrapingBee fetches the page first. The generated extractor does not receive the ScrapingBee API key and does not make its own network requests. It gets HTML through standard input and produces JSON.
  • The verifier then checks that JSON against an independent reference that is not included in the model prompt.

At a high level, the runner does something like this:

code = generate_extractor(html, contract)

stdout, stderr, returncode = execute_extractor(
    code,
    html,
    source_url,
)

result = verify_stdout(
    stdout,
    contract,
    reference,
)

if not result["passed"]:
    code = refine_extractor(
        previous_code=code,
        stderr=stderr,
        verification=result,
        html=html,
    )

The real runner also saves each generated script, stdout, stderr, verification results, model metadata, request metadata, and diffs between revisions.

For the complete implementation, see the source code for this tutorial on GitHub.

The baseline was easier than expected

The model received the HTML from the first Books to Scrape catalogue page and generated an extractor. It passed on the first attempt. I did not want to count a script as successful just because it worked on the exact page shown to the model, so I also ran the same generated code against two holdout pages:

  • catalogue page 2, containing another 20 books;
  • the Travel category, containing 11 books and using a different relative URL depth.

The model had not seen either page while generating the extractor.

[execute] attempt=1 target=catalogue_page_1 role=generation
[verify] PASS target=catalogue_page_1

[execute] attempt=1 target=catalogue_page_2 role=holdout
[verify] PASS target=catalogue_page_2

[execute] attempt=1 target=travel_category role=holdout
[verify] PASS target=travel_category

[verify] PASS attempt=1 targets=3/3

So my first result was not a dramatic self-repair story. The generated code simply worked. That is still a valid outcome. A review loop should not rewrite code just to prove that Refine exists. If Generate → Execute → Verify passes, the correct next step is to keep the verified extractor and move on.

Making the test harder: JavaScript and incomplete pages

The static test gave the generated extractor very little trouble, so I moved the experiment to pages where getting the right HTML is part of the problem. That turned out to be much more interesting.

JavaScript-rendered content

My next target was the JavaScript-rendered products challenge on Scrapify Data Labs. A plain fetch returned a perfectly valid response, but there were no product cards in the HTML:

[fetch/raw] ok status=200 credits=1 product_cards=0

There was nothing a better BeautifulSoup selector could do with that response. The products had not been inserted into the DOM yet.

I fetched the page again with ScrapingBee's JavaScript rendering enabled and waited for the product list to report that it was ready:

[fetch/js] ok status=200 credits=5 product_cards=16 data_ready=True

This time all 16 products were present.

Only then did we generate the extractor:

[generate] attempt 1
[execute] attempt 1
[verify] PASS on attempt 1

The parser again worked on its first attempt. The useful result was elsewhere: the experiment showed that verification sometimes needs to happen before code generation. If the input HTML is incomplete, asking the model to repair extraction code is solving the wrong problem.

Infinite scroll

I then tried the Scrapify Data Labs infinite-scroll challenge, where 48 records appear in batches as the page is scrolled.

A normal JavaScript-rendered request loaded only the first six:

[fetch/initial-js] ok status=200 credits=5 items=6/48 data_complete='false'

The page exposed enough state for the runner to detect that the fetch was incomplete:

[verify/fetch] FETCH_INCOMPLETE items=6/48

No extraction code had been generated yet, so there was nothing to repair.

Instead, I changed the fetch strategy. My first scrolling attempt timed out. I then switched to an explicit sequence that repeatedly scrolled to the bottom of the page, waited for the next batch, and finally waited for the feed to report data-complete="true".

The second strategy returned the complete DOM:

[fetch/refined-js] ok status=200 credits=5 items=48/48 data_complete='true'

At that point, the runner could build its independent reference from the completed page and start code generation.

The first generated script was rejected before execution:

[generate] attempt 1
[verify/code] FAIL on attempt 1
  - Disallowed import: argparse

This was not an extraction failure: my own pre-execution allowlist was too strict and rejected a harmless standard-library import.

After another generation attempt:

[generate] attempt 2
[execute] attempt 2
[verify/code] PASS on attempt 2

An earlier version of the same policy had made a similar mistake with sys.exit().

These runs gave us a more useful experiment than a manufactured broken selector would have. The generated extractors themselves were mostly fine. The harder problems were making sure the page was complete before extraction started and making sure our own checks did not reject valid code.

The next section looks at those failure modes separately, without assuming every bad result means "the LLM wrote the parser wrong."

What actually broke, and how the loop caught it

My experiments did not produce the tidy sequence I expected. The generated extractors themselves worked surprisingly well. The more interesting failures happened around them.

Incomplete input

The infinite-scroll test exposed an interesting example: the first rendered page contained only 6 of 48 records. Rewriting the extractor would have achieved nothing because the other 42 records were not in the HTML yet. This is a separate failure class from bad extraction code. JavaScript rendering, lazy loading, pagination, and infinite scroll can all produce incomplete input while still returning a perfectly successful HTTP response.

That is why verification should start before the parser runs. First make sure the page you are about to parse is actually complete.

My Books to Scrape test did not cover automatic pagination discovery either. I tested the same extractor on additional pages, but following pagination would be a separate behavior to verify.

Plausible output can still be wrong

An extractor can also run successfully and return valid JSON while still extracting the wrong data. A brittle selector may silently miss records. A parser may return the correct number of items but get a price, title, or rating wrong.

I did not hit a genuine value-corruption bug in the final runs, but my verifier was designed for this case. It compared the result with an independent reference rather than asking whether the output merely looked plausible. Syntax validity is therefore a weak success criterion. What matters is whether the extracted data satisfies the checks you actually care about.

Adaptive Python scraping with Scrapling tackles changing page structure from another direction, but adaptive extraction still benefits from independent verification.

Your verifier can fail too

Finally, the review machinery is software, and software can be wrong. My pre-execution policy rejected harmless uses of sys.exit() and argparse. Those were false failures caused by an overly strict allowlist, not bad extraction code. This is another reason not to let the model simply grade its own answer. Verification should be grounded in evidence from real execution: record counts, schema checks, known values, page state, exit codes, or another independent oracle.

And the loop still needs an escape hatch. Two or three Refine attempts is a reasonable default. If the problem is not converging after that, stop, preserve the evidence, and send it for human review.

The main lesson is simple: before refining anything, identify what actually failed — the input, the extractor, or the checks around it.

How this differs from ScrapingBee's ai_query and ai_extract_rules

ScrapingBee already has AI-powered extraction, but it works differently from the code-generation loop in this article.

With ai_query, you describe the information you want in natural language:

ai_query="price of the product"

ai_extract_rules is useful when you want several structured fields. You provide a JSON object whose keys define the output and whose values describe what should be extracted:

{
  "title": "product title",
  "price": "current product price",
  "in_stock": "whether the product is available"
}

ScrapingBee interprets those instructions during the request and returns the extracted result. With ai_extract_rules, the response follows the structure you defined. The full syntax is covered in the data extraction rules documentation.

The important difference is that neither feature produces source code. There is no Python extractor to save, execute separately, verify, or revise. The AI interpretation happens as part of each API request.

A self-reviewing code-generation loop does the opposite: it uses the LLM to produce a reusable program. Once that program has passed verification, you can keep running the same code without invoking the model for every page. Therefore:

  • For a one-off extraction, a small batch, or pages whose structure varies heavily, ai_query or ai_extract_rules may be simpler. You describe what you need and get the data back without maintaining parser code.
  • A generated extractor becomes more interesting when you repeatedly process the same kind of page. The initial generation and verification require more machinery, but afterward the extraction itself can be a small deterministic script that is easy to rerun and inspect.

There is a cost difference as well: both ai_query and ai_extract_rules add 5 credits to the normal request cost. For example, a classic request with JavaScript rendering costs 5 credits, or 10 credits when one of those AI extraction features is added.

The trade-off is that the code-generation approach moves more responsibility to your side. You need an execution environment, verification rules, retry handling, logs, and a safe way to run generated code.

So these are not competing modes of the same ScrapingBee feature. ai_query and ai_extract_rules are single-request extraction features. The loop in this article is application logic you build yourself, with ScrapingBee optionally handling the page-fetching part.

Where a fetch layer like ScrapingBee's MCP server fits into the loop

My test runner calls ScrapingBee's HTTP API directly. If the same workflow were controlled by an AI agent, ScrapingBee's remote MCP server could provide the fetch layer as a set of tools instead.

What the MCP server provides

The current MCP server includes several tools relevant to web extraction:

  • get_page_html — fetch the full page HTML;
  • get_page_text — return the page as text or Markdown;
  • extract_page_data — extract structured data using CSS or XPath rules;
  • get_screenshot — capture a page or element screenshot;
  • get_file — fetch files such as images or PDFs.

For my kind of loop, get_page_html is the most obvious integration point. It can handle JavaScript rendering and browser controls such as waits and js_scenario, along with proxy and geo-targeting options. An agent can therefore fetch the page through MCP and pass the returned HTML into its own Generate → Execute → Verify → Refine workflow.

The important word there is own. The MCP server is not the loop. It does not write an extractor, execute generated Python, decide whether its output is correct, or repair the program after a failed check.

extract_page_data does not change that distinction either. It can return structured data directly from a page using CSS or XPath extraction rules, but it does not produce a reusable source-code artifact and run it through a self-review cycle.

Why keep fetching separate?

My JavaScript and infinite-scroll experiments showed why this boundary is useful. Sometimes the extractor is fine and the input page is not. If JavaScript has not finished loading or an infinite-scroll feed is incomplete, changing the parser will not recover data that never reached the DOM.

An agent can recognize that kind of failure and fetch the page again with different rendering, waits, or a JavaScript scenario. Only after the HTML is complete does it make sense to generate or repair extraction code.

Keeping those responsibilities separate also means the generated scraper does not have to manage browser execution, proxy configuration, geo-targeting, or page interaction itself. That infrastructure stays in the fetch layer, while your application owns code generation and verification.

The same separation is useful if you use ScrapingBee's AI web scraping API for direct extraction instead: fetching and rendering are ScrapingBee's job, while any external code-generation loop remains your own application logic.

For repeated attempts, the agent should still use normal retry discipline: back off after temporary failures and limit how many times it can refetch or revise a result. So MCP's role here is fairly simple: it gives an agent a convenient way to obtain the page it needs. It does not turn ScrapingBee into a self-reviewing code-generation system.

When to use a code-generation loop, and when a single call is enough

A code-generation loop earns its keep when you expect to reuse the result. If you need data from one page once, building generation, execution, and verification infrastructure is probably overkill. A single ai_query or ai_extract_rules request is much simpler.

The balance starts to shift when you repeatedly extract the same kind of page and its structure is reasonably stable. In that case, spending some effort to generate and verify a reusable parser can make sense.

A simple rule of thumb

SituationBetter fit
One-off or small extraction jobai_query / ai_extract_rules
Pages with highly variable structureai_query / ai_extract_rules
Repeated extraction from the same page familyGenerated extractor
Stable or semi-stable markup at meaningful volumeGenerated extractor

Cost can matter here as well. A verified parser can keep running without invoking an LLM for every page, while AI extraction adds work to every request. The exact balance depends on the target, rendering requirements, and request volume, so check the current ScrapingBee pricing rather than assuming one approach is always cheaper.

Remember what you have to own

The reusable-script approach comes with more engineering. You need a controlled execution environment, meaningful verification checks, logging, failure reporting, and a way to inspect cases the loop cannot resolve. The verifier also needs to be good enough to recognize when the system is fixing the wrong thing.

Human review still has a place here. A model may be able to repair a local selector or parsing bug when given useful feedback, but it is much less likely to recover cleanly when the underlying extraction strategy or verification assumptions are wrong.

One scope boundary stays the same whichever approach you choose. ScrapingBee's HTML API only allows pre-login scraping of publicly accessible content; post-login scraping is not allowed. A generated extractor does not change that policy. So the decision is less about whether AI-generated code is "better" than a single extraction call and more about reuse: if you expect the same verified code to run often enough, the extra machinery can be worth owning.

Give your extraction loop reliable pages to work with, with ScrapingBee

A self-reviewing code-generation loop is something you build yourself: you own the generated code, the execution environment, the verification rules, and the retry logic.

ScrapingBee can take care of a different part of the job: fetching the page your loop needs to work with, including JavaScript-rendered pages and more involved browser scenarios like the ones we used in our tests.

New accounts include 1,000 free API credits, and no credit card is required. Try ScrapingBee today

Still have questions about how the loop works or where it fits? The FAQ below covers the common ones.

Self-reviewing code generation FAQs

What is a self-reviewing code generation loop for web extraction?

It is a workflow where an LLM writes extraction code, the code runs against real page HTML, its output is checked, and failed attempts can be revised using execution and verification feedback.

The important part is that the generated code has to prove itself through execution rather than simply looking correct.

Is a self-reviewing code generation loop something ScrapingBee does automatically?

No. It is an application pattern that you build yourself.

ScrapingBee can provide the page HTML, including rendered HTML for JavaScript-heavy pages, but it does not generate, execute, verify, or repair extraction code as part of a built-in self-review loop.

How many Refine iterations should the loop allow before giving up?

Two or three repair attempts is a reasonable starting point.

If the result still fails after that, stop the loop, preserve the execution and verification evidence, and surface the case for human review rather than continuing indefinitely.

Does the loop have to be written in Python?

No. We used Python and BeautifulSoup for the experiments in this article, but the pattern is language-independent.

You can build the same Generate → Execute → Verify → Refine workflow in any language where you can run generated code, parse the target content, and check the result.

Can a self-reviewing code generation loop scrape pages that require a login?

The pattern itself is not what sets that boundary, but the ScrapingBee setup used in this article does.

ScrapingBee's HTML API is limited to publicly accessible, pre-login content, so the examples here do not target pages or data that require authentication. Adding a code-generation loop does not change that policy.

How is this different from a self-healing scraper?

A self-healing scraper usually tries to keep an existing extraction strategy working when page structure changes, for example by locating an element again after a selector breaks.

A self-reviewing code-generation loop does something different: it generates source code, executes it, verifies the result, and can produce a revised version of that code after a failure.

image description
Ilya Krukowski

Ilya is an IT tutor and author, web developer, and ex-Microsoft/Cisco specialist. His primary programming languages are Ruby, JavaScript, Python, and Elixir. He enjoys coding, teaching people and learning new things. In his free time he writes educational posts, participates in OpenSource projects, tweets, goes in for sports and plays music.

Auto-mode picks the configuration that successfully scrapes your page

Try it now