# mcp-page-finder-extractor [Health: Active]

**Category:** 💻 Developer Tools  
**Repository:** https://github.com/mambalabsdev/mcp-page-finder-extractor  
**GitHub Stars:** 0  
**npm Downloads (last month):** 174  
**Views:** 0  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/mcp-page-finder-extractor

## Description
Find a named page type on a company's own website from its domain, in 11 languages.

## Tools
Capabilities this server exposes over MCP:

- **find_company_page**

## Claude Desktop Quick Installation
Install path detected from listing signals. Uses `npx` (confidence: high):

```json
"mcpServers": {
  "mcp-page-finder-extractor": {
    "command": "npx",
    "args": ["-y","@mambalabsdev/mcp-page-finder-extractor"],
    "env": {
      "APIFY_TOKEN": ""
    }
  }
}
```

**Requires environment variables:** `APIFY_TOKEN` — the values above are empty placeholders; fill in real credentials before running (see the repository for what each one is for).

## Documentation

## What mcp-page-finder-extractor does

The mcp-page-finder-extractor MCP server connects an MCP client to the Mamba Labs Page Finder and Extractor actor on Apify. It accepts one domain, several domains, or company records containing fields such as name, country, ISIN, ticker, or URL. The `find_company_page` tool then searches the company’s own public website for selected page types.

The available page types include pricing, free trials, demo requests, investor relations, annual reports, security and compliance pages, privacy and legal pages, careers and job boards, documentation, API references, integrations, blogs, press rooms, case studies, events, marketplaces, support centers, login pages, sustainability, diversity, and more. There are 46 types in total. Discovery vocabulary covers 11 European languages.

## How it works

Use `mode: "locate"` to receive page URLs and discovery metadata, or `mode: "locate_and_extract"` to read the located pages and produce structured fields. Results contain one flat row per input. For each requested type, the output can include the URL, whether the page was found, the discovery method, and a confidence score associated with that method.

A `false` result means the site was read and its links, sitemap, and related paths did not reveal the page. A `null` result means the server could not read enough to determine the answer. Check `coverage` and `fetch_status` before treating a negative result as conclusive. Extracted fields are returned in a separate findings dataset keyed by the input.

## Setup and configuration

Run the package with:

```bash
npx -y @mambalabsdev/mcp-page-finder-extractor
```

Set `APIFY_TOKEN` to an Apify API token. The server starts an Apify actor run and returns its dataset. The mcp-page-finder-extractor MCP server can be configured as a standard stdio server; the README provides a Claude Desktop configuration using `npx`, the package name, and the token in the client environment.

Optional inputs include `pageTypes`, `knownUrls`, extraction field selections, language hints, concurrency, request ceilings, candidate-page limits, browser rendering, and cache control. `knownUrls` can bypass discovery for page types whose URLs are already available. Setting `skipCache` to `"true"` requests a fresh crawl instead of using the 14-day cache.

## Tools and capabilities

- `find_company_page`: locate or extract requested page types for domains, domain lists, or company identities.
- Return the method used, including known URL, homepage or footer anchor, section hop, sitemap, or path guess.
- Use browser rendering for publicly served pages that require JavaScript when `allowRender` permits it.
- Process multiple companies concurrently while retaining one request at a time per company.
- Supply page-type-specific and page-agnostic extraction fields.

## Limitations and notes

The mcp-page-finder-extractor MCP server is read-only. It reads publicly available pages on the company’s own website and writes nothing elsewhere. It follows each host’s `robots.txt` rules and crawl delay, identifies itself with a descriptive user agent, and does not sign headers, impersonate a browser, or retry around access blocks. Browser rendering is not used against access controls.

Usage is billed through Apify credits. Locate, extraction, and company-name identity resolution have separate event charges; identity resolution is only used with the `companies` input. Completed looks are billed even when no page is found, while dead domains, refusals, and robots exclusions return `found: null` without a locate charge.

The `contact` and `about` page types may return personal data that the company publishes on those pages. Such records include `is_personal_data` and a `lawful_basis`. The tool does not infer attributes or perform outside lookups, and it is not intended as a general employee-roster extractor.

_Full upstream README: https://allmcps.com/mcp/mcp-page-finder-extractor/readme_

