---
title: "How to Turn Any Website Into Clean Markdown for ChatGPT, RAG or a Custom GPT"
slug: "turn-website-into-llm-knowledge-base"
type: "blog-post"
url: "https://automationbyexperts.com/blog/turn-website-into-llm-knowledge-base"
date: "2026-10-02"
category: "AI"
tags: ["RAG", "LLM", "Web Crawling", "Markdown", "Custom GPT", "Apify"]
read_time: 9
author: "Youssef Farhan"
---

# How to Turn Any Website Into Clean Markdown for ChatGPT, RAG or a Custom GPT

> Crawl a whole website or docs site into clean Markdown with the navigation and ads stripped out, ready to upload to a Custom GPT, a Claude Project or a vector database. Which Apify tool to use, how to configure it and what it costs.

**Short answer:** to feed a whole website to an AI, crawl it with [Website Content Crawler](https://apify.com/apify/website-content-crawler?fpr=youssef&fp_sid=abe_blog_rag), which follows the site's links, strips menus, footers and cookie banners, and saves each page as clean Markdown. Download the result as files for a Custom GPT or Claude Project, or push it into a vector database for a RAG pipeline. If you instead need fresh answers from Google at question time, use [RAG Web Browser](https://apify.com/apify/rag-web-browser?fpr=youssef&fp_sid=abe_blog_rag), which searches and returns the top pages as Markdown.

Both tools are built and maintained by Apify itself, not by me. They are among the most used actors on the platform, with around 160,000 and 180,000 users respectively, and they are on my shortlist of [the best Apify actors](https://automationbyexperts.com/best-apify-actors) for jobs outside my own catalogue. This guide explains when to use which, the settings that decide whether your AI gets clean text or junk, and what it costs.

## Why not just give ChatGPT the URL?

Because a chat model browsing a URL reads one page at a time, often only partly, and forgets it at the end of the conversation. A knowledge base needs the whole site, stored once, in a format the model reads well. Raw HTML is a poor format: most of a typical page is navigation, scripts, cookie notices and footer links, and every one of those wastes tokens and confuses retrieval. Markdown keeps what matters to a model (headings, lists, tables, links) and drops the rest.

## Which tool should you use?

- **Website Content Crawler** when you know the site: your docs, your help centre, a competitor's product pages, a regulator's guidance pages. You give it start URLs and it crawls everything under them.

- **RAG Web Browser** when you know the question but not the site. It runs a Google search, opens the top results in a real browser and returns their text as Markdown. It is designed to be called by an AI agent at question time, and can run in a fast standby mode for low latency. Note that its Google search runs from the United States.

A simple rule: stored knowledge you will reuse, crawl it; one-off current questions, search it.

## How do you crawl a website into Markdown?

- **Pick the start URL carefully.** The crawler stays under the start URL's path, so `https://example.com/docs/` crawls the docs and skips the blog and marketing pages. That one choice removes most of the noise.

- **Choose the crawler type.** A plain HTTP crawler is fast and cheap and works for static sites and most documentation. Use the headless browser option only for sites that build their content with JavaScript. If the first test returns empty pages, switch to the browser.

- **Set limits.** Set a maximum number of pages and a maximum crawl depth for the first run. Check what came back before crawling ten thousand pages.

- **Keep the cleaning on.** The default settings remove navigation, headers, footers, cookie banners and other boilerplate. If a site uses unusual markup, add CSS selectors for elements to remove.

- **Choose the output.** Markdown is the right default for AI use. You can also save plain text or HTML, and download linked files such as PDFs.

- **Export.** Each page is one row with its URL, title, metadata and Markdown. Download the dataset, or use the built-in integrations to send it straight to a vector database.

## How do you load the result into a Custom GPT or Claude Project?

Custom GPTs and Claude Projects accept uploaded files but have limits on the number and size of files. Hundreds of tiny page files are awkward, so combine pages into a few large Markdown files grouped by section, for example one file for "getting started" pages and one for "API reference". Keep each page's URL as a heading line above its text: when the model answers, it can then tell you which page it used, and you can check it.

For a bigger or frequently changing site, a proper RAG setup is better: split pages into chunks, embed them, and store them in a vector database such as Pinecone or Qdrant. Website Content Crawler integrates directly with the common vector databases and with LangChain and LlamaIndex, so the crawl can feed your index without glue code.

## How much does it cost?

RAG Web Browser uses per-event pricing. At the time of writing it charges $0.00348 per search result and $0.00228 per fetched page, plus a tiny run start fee. A question that searches Google and reads the top 3 pages costs about 2 cents.

Website Content Crawler is billed on Apify platform usage, meaning the compute time the crawl takes, rather than a fixed price per page. That makes the crawler type the main cost lever: the plain HTTP crawler processes pages far faster, and therefore more cheaply, than a headless browser. Crawl a small sample, look at the usage figure for the run, and scale from there. Apify's free plan includes $5 of monthly credit, which is plenty for trying either tool on a small site.

## What goes wrong, and how do you avoid it?

- **Crawling the whole domain.** Start at the section you need. A docs crawl that wanders into the blog and changelog fills your index with irrelevant text.

- **Duplicate pages.** Printable versions, tag pages and URLs with tracking parameters produce duplicates. Exclude those URL patterns.

- **Stale knowledge.** Sites change. Schedule the crawl weekly or monthly and rebuild the index, or your assistant quotes last year's pricing with confidence.

- **Other people's content.** Crawling your own site is simple. For third-party sites, respect robots.txt and terms, and use the content for internal research rather than republishing it.

## What can you do with the Markdown next?

Once a site is in clean text, you can run questions over every page. The [Bulk LLM Runner](https://automationbyexperts.com/apify/bulk-llm-runner) runs one prompt per row, so a crawl of a competitor's product pages can become a table of product name, price, target customer and key claims in one pass. If the goal is making your own site easier for AI to read, the [llms.txt generator](https://automationbyexperts.com/apify/llms-txt-generator) builds the index file that AI crawlers look for, and the [SEO, GEO and AEO audit](https://automationbyexperts.com/apify/seo-geo-aeo-audit) checks how answer engines see your pages.

## Frequently asked questions

**What is the best way to turn a website into Markdown for an LLM?**Crawl it with a crawler that removes boilerplate and outputs Markdown, such as Apify's Website Content Crawler. Start from the section you need, test on a few pages, then crawl the rest and export one Markdown document per page.

**What is the difference between Website Content Crawler and RAG Web Browser?**Website Content Crawler crawls sites you choose and stores their pages, which suits a reusable knowledge base. RAG Web Browser searches Google for a query and returns the top pages as Markdown on demand, which suits AI agents answering current questions.

**Can I use the crawl for a Custom GPT?**Yes. Combine the page Markdown into a few larger files grouped by topic, keep each page URL as a heading, and upload them as the GPT's knowledge. The same files work for Claude Projects.

**Does it work on JavaScript-heavy sites?**Yes. Switch the crawler type to a headless browser, which renders JavaScript before extracting text. It is slower and costs more than the plain HTTP crawler, so use it only where needed.

**Is it legal to crawl a website for AI?**Crawling your own content is fine. For other sites, follow robots.txt and the site's terms, avoid personal data, and use the text for internal analysis rather than republishing it.

Need data from a specific site instead of its text? My [catalogue of scrapers](https://automationbyexperts.com/apify) returns structured fields for sites like AutoTrader, Wallapop and Fnac, and you can [request a new one](https://automationbyexperts.com/suggest).

## Actors used in this guide

- [Run GPT, Claude, DeepSeek... Prompts in Bulk (No API Key)](https://automationbyexperts.com/apify/bulk-llm-runner): It runs hundreds of prompts in parallel across GPT, Claude, Gemini, Perplexity, DeepSeek, Qwen, Kimi and 350+ other m…
- [LLMs.txt & llms-full.txt Generator for Any Website](https://automationbyexperts.com/apify/llms-txt-generator): It generates a spec-compliant llms.txt and llms-full.txt for any website in one click - crawling your sitemap or site…
- [SEO, GEO & AEO Audit - AI Search Readiness Checker](https://automationbyexperts.com/apify/seo-geo-aeo-audit): It audits any website for Google SEO and AI-search visibility in one run - returning 0–100 SEO, GEO and AEO scores pe…
