Source-capture policy
What our crawlers do, and what they never do. If you have seen MINT-Capture in your server logs, or MINT-Spine in the FCA's, this is the policy they run under and the address to write to.
How to identify us, and how to reach us
Every request we make carries a User-Agent naming a product token, a link to this page, and a contact email address; the same address is repeated in a From: header. There are two tokens, because there are two jobs:
- MINT-Capture reads firms’ own websites for published fee information. It is the one that reads
robots.txt, and the one the rest of this page is about. - MINT-Spine calls the FCA’s documented Register API and nothing else. It never touches a firm’s website, so
robots.txtdoes not arise; it is listed here because its requests link to this page and you may see it named in the FCA’s logs rather than your own.
Either crawler refuses to start at all if no contact address is configured, so a request from us that you cannot reply to is a request we did not make. If you would rather we did not read your site, say so in your robots.txt or in your website terms and we will stop and not come back — see “How to make us stop” below. You can also write to us at [email protected], and you can use the complaints and disputes form for anything you want removed or corrected.
If you would rather we did not read your site, say so in your robots.txt or in your website terms and we will stop and not come back — see "How to make us stop" below. You can also write to us at [email protected], and you can use the complaints and disputes form for anything you want removed or corrected.
What we read
- Publicly served pages and PDFs on a financial-advice firm's own website that state what the firm charges — typically a fees, charges or pricing page.
- At most 15 pages from any one site in a discovery run, and never more than two links deep from the home page.
- Never a page behind a login, a form, a client portal, an account area, a cart or a checkout, and never a
mailto:,tel:orjavascript:link. - Never a page on another organisation's domain reached from yours.
How politely we read it
- We fetch and read your robots.txt before anything else on your site, through the same identified crawler, and we keep a copy of it as the record of what we were told.
- We leave at least 5 seconds between requests to the same site, however many jobs are running. A longer
Crawl-delayin your robots.txt is honoured. ACrawl-delayabove 120 seconds is treated as "do not crawl", and we stop. - We check your website terms before the first content request. If a sentence prohibits automated access, we stop and record the sentence. Only a person here can decide that a flagged term does not apply, and the record shows who decided and when; software never decides that on its own.
- We follow at most five redirects, we record every hop, and we re-run all of these checks at every hop.
- We time out after 30 seconds and do not store a response body over 10 MB.
How to make us stop
Any one of the following stops us permanently for your whole site. There is no expiry, no retry and no appeal by us: once your site is on the stop list, nothing in the software takes it off, and we will not come back with a different User-Agent, a headless browser or a proxy.
- A
Disallowin robots.txt that covers the page, forMINT-Captureor for*. - A 401, 403 or 429 response, from your site or from your robots.txt.
- A bot-management challenge of any kind — including a challenge served with a 200 status.
- A
noaiornoimageairobots directive, in a meta tag or anX-Robots-Tagheader. - A text-and-data-mining reservation (
tdm-reservation: 1), in a header or a meta tag. - A sentence in your terms prohibiting automated access, crawling, scraping, data mining or systematic extraction.
- A
Crawl-delaylonger than 120 seconds.
What we store, and where it goes
- The bytes we received are stored privately, addressed by their own fingerprint, written once and never overwritten or edited. They are evidence of what your page said on the day we read it. They are never republished, re-hosted or served to anyone.
- Beside them we store the record of the fetch: the date and time, the exact User-Agent, the address requested and the address finally served, every redirect, the HTTP status, the robots.txt decision, the terms check, and the version of this policy in force.
- We publish the facts we extract — the charges themselves — with the address they came from, the date we captured them and a link to dispute them. Where we quote your wording it is a short acknowledged excerpt, never a reproduction of the page.
- A PDF is stored as we received it and is never republished.
What we do with the figures
Extracted figures become a declared — published record on your firm's page, with the source link and the capture date beside them and a link to correct or dispute them. Nothing is published from a capture on the software's own confidence alone: a person or a reviewer records a verdict on every extracted schedule first, and the extractor's confidence is capped below the level that would allow anything to publish automatically.
The same figures contribute to the UK Advice Fee Index, in aggregate, under the rules on the methodology hub. The current benchmark page is at advice fees.
What we keep as a record for you
If you ask what our crawler did on your site, we can show you: the robots.txt we read and its fingerprint, the terms check and any sentence it flagged, any decision a person made about it and who made it, every block and the signal that caused it, the full fetch log, and the record of every page we stored. Ask through the complaints and disputes form.
Copyright
Your fee page is your copyright work. We do not republish it. We publish facts extracted from it — figures are not protected by copyright — with attribution, a link and the capture date, and any quotation is short, acknowledged and limited to what is needed to show where a figure came from. If you believe we have gone further than that, tell us and we will take it down while we check.
About this page: information only, from a publisher that is not authorised or regulated by the FCA, sells no advice and endorses no firm. See how we make money, how we rank and the methodology hub. How we handle personal data is set out in the privacy notice.