Short answer: to feed a whole website to an AI, crawl it with Website Content Crawler, which follows the site's links, strips menus, footers and cookie banners, and saves each page as clean Markdown. Download the result as files for a Custom GPT or Claude Project, or push it into a vector database for a RAG pipeline. If you instead need fresh answers from Google at question time, use RAG Web Browser, which searches and returns the top pages as Markdown.

Both tools are built and maintained by Apify itself, not by me. They are among the most used actors on the platform, with around 160,000 and 180,000 users respectively, and they are on my shortlist of the best Apify actors for jobs outside my own catalogue. This guide explains when to use which, the settings that decide whether your AI gets clean text or junk, and what it costs.

Why not just give ChatGPT the URL?

Because a chat model browsing a URL reads one page at a time, often only partly, and forgets it at the end of the conversation. A knowledge base needs the whole site, stored once, in a format the model reads well. Raw HTML is a poor format: most of a typical page is navigation, scripts, cookie notices and footer links, and every one of those wastes tokens and confuses retrieval. Markdown keeps what matters to a model (headings, lists, tables, links) and drops the rest.

Which tool should you use?

  • Website Content Crawler when you know the site: your docs, your help centre, a competitor's product pages, a regulator's guidance pages. You give it start URLs and it crawls everything under them.
  • RAG Web Browser when you know the question but not the site. It runs a Google search, opens the top results in a real browser and returns their text as Markdown. It is designed to be called by an AI agent at question time, and can run in a fast standby mode for low latency. Note that its Google search runs from the United States.

A simple rule: stored knowledge you will reuse, crawl it; one-off current questions, search it.

How do you crawl a website into Markdown?

  1. Pick the start URL carefully. The crawler stays under the start URL's path, so https://example.com/docs/ crawls the docs and skips the blog and marketing pages. That one choice removes most of the noise.
  2. Choose the crawler type. A plain HTTP crawler is fast and cheap and works for static sites and most documentation. Use the headless browser option only for sites that build their content with JavaScript. If the first test returns empty pages, switch to the browser.
  3. Set limits. Set a maximum number of pages and a maximum crawl depth for the first run. Check what came back before crawling ten thousand pages.
  4. Keep the cleaning on. The default settings remove navigation, headers, footers, cookie banners and other boilerplate. If a site uses unusual markup, add CSS selectors for elements to remove.
  5. Choose the output. Markdown is the right default for AI use. You can also save plain text or HTML, and download linked files such as PDFs.
  6. Export. Each page is one row with its URL, title, metadata and Markdown. Download the dataset, or use the built-in integrations to send it straight to a vector database.

How do you load the result into a Custom GPT or Claude Project?

Custom GPTs and Claude Projects accept uploaded files but have limits on the number and size of files. Hundreds of tiny page files are awkward, so combine pages into a few large Markdown files grouped by section, for example one file for "getting started" pages and one for "API reference". Keep each page's URL as a heading line above its text: when the model answers, it can then tell you which page it used, and you can check it.

For a bigger or frequently changing site, a proper RAG setup is better: split pages into chunks, embed them, and store them in a vector database such as Pinecone or Qdrant. Website Content Crawler integrates directly with the common vector databases and with LangChain and LlamaIndex, so the crawl can feed your index without glue code.

How much does it cost?

RAG Web Browser uses per-event pricing. At the time of writing it charges $0.00348 per search result and $0.00228 per fetched page, plus a tiny run start fee. A question that searches Google and reads the top 3 pages costs about 2 cents.

Website Content Crawler is billed on Apify platform usage, meaning the compute time the crawl takes, rather than a fixed price per page. That makes the crawler type the main cost lever: the plain HTTP crawler processes pages far faster, and therefore more cheaply, than a headless browser. Crawl a small sample, look at the usage figure for the run, and scale from there. Apify's free plan includes $5 of monthly credit, which is plenty for trying either tool on a small site.

What goes wrong, and how do you avoid it?

  • Crawling the whole domain. Start at the section you need. A docs crawl that wanders into the blog and changelog fills your index with irrelevant text.
  • Duplicate pages. Printable versions, tag pages and URLs with tracking parameters produce duplicates. Exclude those URL patterns.
  • Stale knowledge. Sites change. Schedule the crawl weekly or monthly and rebuild the index, or your assistant quotes last year's pricing with confidence.
  • Other people's content. Crawling your own site is simple. For third-party sites, respect robots.txt and terms, and use the content for internal research rather than republishing it.

What can you do with the Markdown next?

Once a site is in clean text, you can run questions over every page. The Bulk LLM Runner runs one prompt per row, so a crawl of a competitor's product pages can become a table of product name, price, target customer and key claims in one pass. If the goal is making your own site easier for AI to read, the llms.txt generator builds the index file that AI crawlers look for, and the SEO, GEO and AEO audit checks how answer engines see your pages.

Frequently asked questions

What is the best way to turn a website into Markdown for an LLM?

Crawl it with a crawler that removes boilerplate and outputs Markdown, such as Apify's Website Content Crawler. Start from the section you need, test on a few pages, then crawl the rest and export one Markdown document per page.

What is the difference between Website Content Crawler and RAG Web Browser?

Website Content Crawler crawls sites you choose and stores their pages, which suits a reusable knowledge base. RAG Web Browser searches Google for a query and returns the top pages as Markdown on demand, which suits AI agents answering current questions.

Can I use the crawl for a Custom GPT?

Yes. Combine the page Markdown into a few larger files grouped by topic, keep each page URL as a heading, and upload them as the GPT's knowledge. The same files work for Claude Projects.

Does it work on JavaScript-heavy sites?

Yes. Switch the crawler type to a headless browser, which renders JavaScript before extracting text. It is slower and costs more than the plain HTTP crawler, so use it only where needed.

Is it legal to crawl a website for AI?

Crawling your own content is fine. For other sites, follow robots.txt and the site's terms, avoid personal data, and use the text for internal analysis rather than republishing it.

Need data from a specific site instead of its text? My catalogue of scrapers returns structured fields for sites like AutoTrader, Wallapop and Fnac, and you can request a new one.

Need help implementing this?

I build custom automation, scraping pipelines, and AI solutions for businesses. 170+ projects delivered with a 5.0 rating across 172 reviews. Tell me about your project - I reply within 24 hours.

Start Your Project →