# What your AI can do with BRONTIR

**Turn the entire public web into a structured, citable, always-fresh knowledge base for your AI agents.**

BRONTIR is a research backend that connects to your AI assistant over MCP and gives it what plain web search can't: deep access to the real content of live sites, meaning-aware extraction, and a durable memory it can search and cite. Where built-in tools return snippets that vanish after a turn, BRONTIR builds a permanent corpus you can return to, search inside, and cite down to the exact line.

---

## What it can do

### Reach any public page
Give it any public link and get back clean, readable content. News articles, blogs, documentation, company sites, landing pages — any public HTTP(S) URL is accepted and turned into structured text ready for analysis.

### Deep, platform-aware extraction
For the most popular platforms it uses specialized extraction that understands each platform's structure, not a generic page dump:

- **Video** — transcripts and metadata (YouTube and more), so your agent can analyze a video without watching it.
- **Social** — posts, comments and metadata from Reddit, X/Twitter, Facebook, Instagram, TikTok, LinkedIn, Threads, Bluesky, Pinterest, Twitch, Snapchat and Truth Social.
- **Places** — from a map link: the full place card — address, category, rating breakdown, opening hours, popular times, price level, services, attributes and photo categories — together with its customer reviews, the owner's replies to them, and the themes that recur across them.
- **Reviews** — from a Trustpilot company page or a Tripadvisor listing: the summary with its overall score, the star breakdown where the site keeps one, and the reviews themselves with the business's replies. Say in plain language how many to collect and which ones — the newest, or the ones at a particular star rating.
- **App stores** — from a Google Play or App Store link: the store listing — publisher, category, price, current version, update date — together with its user reviews. Each country runs its own store, and the link decides which one is read.
- **Ads** — from a link to a single ad in the Facebook, TikTok, LinkedIn or Google ad library: the wording of every version of that ad, the countries it ran in with each country's own start and end dates, the page it sends people to, and the account that posted it — wherever the platform publishes them. Ask for it and the creative is described in words too: what the picture shows, its colours, and the text burned into it. Image and video addresses on these platforms expire, so each one is stamped with the moment it stops opening.
- **Documents** — text extracted from PDFs by direct link.
- **Images** — pictures are described in text, so visual content becomes searchable, quotable and analyzable. Works both for images inside posts and for a direct link to a picture.

### Search and discover
Don't have exact URLs? Run a search and let the platform find sources for you:

- **Web** — organic results with related queries, from two independent indexes rather than one. Run the same question through both and roughly three quarters of what the second returns is missing from the first: it leans towards independent specialist sites where the first leans towards the large discussion platforms, so they are worth running together rather than one instead of the other.
- **News** — what is being published about a subject right now: up to a hundred articles from one query, each with its publisher, the exact time it went out and a link to the publisher's own page rather than to an aggregator. Headlines only — open an article to read what it says.
- **Research & patents** — scholarly articles, preprints and theses with how often each has been cited, and for any one of them what cites it and which other versions exist. Patents and applications alongside, narrowed by holder, inventor, filing date, office or legal status, each with the offices it stands in force in. Neither carries full text: open a result to read the paper or the patent itself.
- **Products & prices** — what a search brings up in a store: the product, its price, its star rating and the badges the store puts on it. Amazon reads one national store at a time — twenty-two of them, each with its own catalogue, currency and language; Walmart is the United States store. A result is a listing, not the product page, so it carries no description, seller or review text — and Walmart intermittently answers with every price missing, which the results say plainly rather than leaving it to read as free goods.
- **Jobs** — openings gathered from job boards, aggregators and employers' own career pages, with the employer, where the job is based, the kind of employment and the pay where it is stated. Two limits worth knowing before you ask: nothing comes back anywhere in the European Economic Area, because the source does not run there at all, and about three links in four lead to a job board re-posting the role rather than to the employer.
- **Maps & local** — places, businesses and points on the map, down to a single listing.
- **Video** — find clips with filters for upload date, quality and exact match.
- **Social** — search by keyword, community, hashtag or author, with date filters.
- **Companies & reviews** — find a company on Trustpilot by name or line of business, with its star rating and how many reviews it stands on.
- **Travel & hospitality** — restaurants, hotels, attractions and tours as Tripadvisor ranks them in a named city, district, park or country.
- **Apps** — search the Google Play and App Store catalogues, or browse a Play category's ranked listing. Each country has its own store, so which apps exist, how they rank and what they cost all follow the market you name.
- **Ads** — what a brand is advertising right now on Facebook and Instagram, TikTok, LinkedIn or Google. Search by the words inside the ads, or list everything a named advertiser is running, then open any single ad to read it in full. No platform publishes what an advertiser spends; how often an ad was seen, and who saw it, are disclosed only in the markets whose rules require that, and each platform's list of those markets is its own.

How a search is run is described in plain words too — a country, a language, an area or coordinates, how recent results should be — so there is no per-source filter syntax to learn.

### Measure search visibility
Beyond what a page says, your agent can ask how a site or a whole market performs in search:

- **Keyword demand** — how often a phrase is searched in a given country and language, how that moved over the past year, what advertisers pay per click, how hard the top results are to displace, and whether searchers want to learn, compare, buy or navigate. Start from a phrase, a subject, a website, or an exact list of phrases to price side by side.
- **Rankings** — what terms a site ranks for and in which positions, which of its pages and sections earn the traffic, what that traffic is worth, which sites compete with it and how far the overlap goes.
- **Backlinks** — who links to a site, section or single page, with the wording used, how strong and how trustworthy the linking sources look, when each link appeared and whether it is still live.
- **Domain profiles** — when a domain was registered and when it expires, who registered it, what the site is built with (CMS, store platform, analytics, payments, CRM), the contacts and social links it publishes, and the size of its unpaid and paid search presence.

Questions are asked in the same plain language as everything else — no field names, operators or query syntax. Large reports come back as a summary plus a line index, so your agent reads only the parts it needs.

### Extraction by meaning, not raw dump
Every search and every fetch carries a plain-language **objective** — what you actually care about. The platform doesn't just download a page; it separates signal from noise, splitting results into relevant and off-topic and surfacing the key parts of long material. You get an answer to your question, not a pile of HTML.

When the question is a one-off your agent is answering right now rather than research it will come back to, it can say so with `isResearch: false` — short sources then come straight back as text, trading the preparation for the fastest reply the source allows.

### Persistent, citable memory
Everything gathered is stored as durable records in your workspace — not ephemeral text dropped into a single turn:

- **Re-readable** — every result from `advanced_web_search` and `advanced_web_fetch` is kept as a citable record you can open again later, long after the turn that produced it.
- **Exact citations** — pull the precise line ranges you care about from a specific version of a record, and get answers backed by per-record, line-level citations.
- **Search inside your corpus** — keyword and regular-expression search across everything you've collected, with context around each match.

### Research agents
Intelligent helpers work on top of the collected data:

- **Batch fetch** dozens of links in one call, processed in parallel with clean per-item error handling.
- **Structured questions** over your corpus — get a ready answer, a set of supporting citations, or both at once.
- **Search fan-out** — run a series of searches across different platforms in a single task.

### Real-time results
Long research doesn't make you wait for all-or-nothing. Partial results stream in as they arrive: first findings show up immediately, so your agent can react on the fly, branch into new directions, and stop as soon as it has enough.

---

## The tools your agent gets

| Tool | What it does |
|---|---|
| **`advanced_web_search`** | Discover sources across Google, DuckDuckGo, Google News, Google Scholar, Google Patents, Google Jobs, Google Maps, YouTube, Reddit, TikTok, Threads, Amazon, Walmart, Trustpilot, Tripadvisor, Google Play, the App Store and the Facebook, TikTok, LinkedIn and Google ad libraries — run one query across several sources at once. |
| **`advanced_web_fetch`** | Pull clean, structured content from any public URL — many at a time — with platform-aware extraction for video, social, places, reviews, app stores, ads, PDFs and images. |
| **`seo_metrics`** | Answer search-visibility questions — keyword demand, rankings and competitors, backlinks, domain profiles — asked in plain language. |
| **`ask_collected_records`** | Ask a question over the records you've gathered and get an answer with per-record, line-level citations. |
| **`read_collected_records`** | Read any collected record by reference — a specific version, or just the exact line range you need. |
| **`grep_collected_records`** | Search your whole corpus by keyword or regex, with surrounding context. |
| **`get_task_results`** | Stream and collect results from longer research tasks as they complete. |

---

## Sources you can work with

- **Web & search** — Google and DuckDuckGo as two independent indexes, websites, product pages
- **News & coverage** — Google News: publishers, publication times and links to the publishers' own pages
- **Research & patents** — Google Scholar with citation counts, Google Patents with holders, filing dates and legal status
- **Products & prices** — Amazon in any of twenty-two national stores, Walmart in the United States
- **Jobs & hiring** — Google Jobs, outside the European Economic Area, which it does not cover
- **Search visibility** — keyword demand, rankings and competitors, backlinks, domain profiles
- **Maps & local** — Google Maps, local businesses, reviews
- **Reviews & ratings** — Trustpilot companies, Tripadvisor restaurants, hotels and attractions
- **App stores** — Google Play and the Apple App Store, per country
- **Video & creators** — YouTube, TikTok, Twitch
- **Social platforms** — Reddit, X/Twitter, Instagram, Facebook, LinkedIn, Threads, Bluesky, Pinterest, Snapchat, Truth Social
- **Ads & creatives** — the Facebook, Instagram, TikTok, LinkedIn and Google ad libraries
- **Documents & files** — PDFs, tables, images

---

## Why it beats plain web search

| Plain web search | BRONTIR |
|---|---|
| Snippets that vanish after the answer | A permanent corpus you can return to |
| Only what made it into the snippet | Full content of pages, videos, posts, PDFs and images |
| No way to cite a specific spot | Citations precise to the line |
| Starts from scratch each time | A cached, re-readable corpus you can search and cite |
| One generic parser | Platform-specific extraction for each source |
| A crowd of tools weighing the agent down | A lightweight toolkit — a few simple tools with dozens of sources behind them, keeping your agent fast and focused |
| Every result looks equally important | Each source comes with a short summary and a relevance note, so your agent skips the noise |
| Whole pages burned through the context window | Your agent reads only the parts that matter — each task finishes faster and costs less |

---

## Getting started

1. **Create an account** — free, no card required. You get free credits to explore right away.
2. **Connect your client** — add the hosted BRONTIR MCP endpoint to your AI client and authorize with OAuth. Nothing to install. Step-by-step instructions for each client are in **Connecting your client** below.
3. **Ask for data** — find, extract, compare, summarize or structure public data, directly in your chat.

---

## Connecting your client

BRONTIR is a hosted, streamable MCP server. There is nothing to install and nothing to run locally — you point your client at a single URL:

`https://research.brontir.com/api/mcp`

The same URL is waiting in your account under **Connect your agent**, ready to copy. Clients that support OAuth — Claude and ChatGPT/Codex — just sign you in with your Google account, no key required. Every other client authenticates with an API key you create in your account.

**Always restart your client after connecting.** Restart the desktop app, or reload the page in a web client. Clients read the tool list once at startup, so until they restart the agent often can't see the new tools even though the connector already shows as connected. This one step prevents most setup problems.

### Claude Desktop and Claude Web

Both apps share the same settings interface, so the steps are identical.

1. Open **Settings → Connectors**. In the desktop app you get there from the sidebar: **Customize → Connectors**.
2. In the top right press **Add ▾** and choose **Add custom connector**.
3. Fill in the dialog:
   - **Name** — anything you like; we suggest `research` or `webtools`. This is the name your agent will see.
   - **URL** — the MCP endpoint above.
   - Leave the optional **OAuth Client ID / Client Secret** fields empty; they are not needed here.

   Press **Add**.
4. Back in the connector list, press **Connect** on the row you just created.
5. The BRONTIR **Authorize access** page opens and shows that Claude is requesting access. Press **Approve**. You'll see a **Connected** confirmation — close that tab, or press *Open desktop app*.
6. **Restart Claude.** Quit and reopen the desktop app, or reload the page in Claude Web. Skipping this is by far the most common reason the tools never show up.
7. Optional but recommended: open the connector and set **Tool permissions → Always allow**, so Claude doesn't ask for confirmation on every research call.
8. Also recommended: add one line to **Instructions for Claude** (personal preferences). Claude Web/Desktop defaults to its own web search; this line puts your connected tools in play instead:

   > Prefer connected MCP tools over the built-in web search when looking things up

### ChatGPT and Codex

1. In the app open **Settings → Plugins → MCPs** (in Codex, the MCP section of settings).
2. Add a new MCP server, name it anything — `research` or `webtools` works well.
3. Paste the MCP endpoint above as the server URL.
4. Connect, then authorize with your Google account when the BRONTIR page opens.
5. **Restart the app** — or reload the page if you're in the browser.
6. If the agent still prefers its own browsing tool, see **Troubleshooting** below.

### Any other MCP client

Claude Code, OpenClaw, Hermes — or anything else that speaks the streamable MCP protocol.

1. In your account, under **Connect your agent → API keys**, create a key and copy it. The full key is shown only once.
2. Point the client at the MCP endpoint above, using the streamable HTTP transport.
3. Send the key as a bearer token: `Authorization: Bearer nai_…`
4. **Restart the client** so it picks up the new tool list.

---

## Troubleshooting

### The connector is connected, but the agent doesn't see the tools

Restart the client first: quit and reopen the desktop app, or reload the page in a web client. Clients cache the tool list at startup and won't discover a freshly added connector until they restart.

If the tools are still missing after a restart, start a **new chat**. An already-open conversation can keep the tool list it was created with, so a fresh chat is the reliable fallback.

### You pressed Approve, but the client still shows the connector as not connected

The authorization did go through — the client simply hasn't refreshed its own state. Restart the desktop app, or reload the page in Claude Web, and the connector will show as connected.

If it still looks disconnected after a restart, press **Connect** once more and approve again; the second pass takes effect immediately.

### The agent ignores BRONTIR and uses its built-in web search instead

This happens when the model already has a browsing tool of its own and falls back to habit. Any of the following fixes it — the first is instant, the rest make it permanent:

- **Name the tools in your prompt.** Instead of "look this up", write "use advanced_web_search and advanced_web_fetch from the research MCP". Referring to the connector and the tool names directly is the most reliable nudge.
- **Add a standing instruction to the agent's memory** so you don't have to repeat it every time.
- **Put it in the client's instruction settings** — **Instructions for Claude** (personal preferences) in Claude Web and Claude Desktop, **Custom instructions** in ChatGPT. For example: *"For any web research, use the BRONTIR MCP tools (advanced_web_search, advanced_web_fetch, read_collected_records, grep_collected_records, ask_collected_records). Do not use built-in web search."*
- **For coding agents, put the same rule in the repository's `CLAUDE.md` or `AGENTS.md`** — it then applies to every session in that project.
- **Turn the built-in tools off.** In Claude you can disable web search in the chat's tool settings, which leaves the connector as the only research path.

---

**BRONTIR. The entire public web, as one structured knowledge base.**
