Web and crawler
FACE reads the web in two ways:
- a web connection reads one configured page as rows of text
- the crawler service scrapes a page, or runs a bounded crawl from a seed URL
Web connection
| Type spellings | webtelemetry (web_telemetry), webpage (the “Webpage” card), web, url, website |
| Settings (properties) | url (also endpoint, target_url, page_url, website, host) |
| Schemes | http and https only |
| Internal addresses | Refused by default. A page named by a URL is normally public. Set allow_internal_network=true to read an intranet page. |
- Execute fetches exactly one page and follows no links:
- 30-second timeout, body capped at 8 MiB.
- At most 5 redirects, on the same host only.
- Only text and HTML are accepted. A document or image is refused, with a pointer to the file path.
- FACE extracts the readable text and splits it into chunks of up to 2,048 characters, with at most 400 chunks.
- Rows have the columns
url,final_url,fetched_at,http_status,content_type,title,chunk_index,chunk_count,chunk_chars,content,truncated,sourceandtransport.truncatedis true when the body or the chunk count hit its cap. - Test fetches the page, reading at most 64 KiB, within 15 s, and applies the same address and redirect rules.
- Explore returns
supported=false. The connector reads one page by design.
The request’s User-Agent is Mozilla/5.0 (compatible; RuninkFACE/1.0) unless
the operator sets SCRAPER_USER_AGENT.
Crawler service
CrawlerService exposes:
| RPC | Behaviour |
|---|---|
Scrape | Reads one url and returns its title, main content and metadata. render_js is reserved. |
Crawl | Starts from url and visits linked pages, returning a crawl_id, a status and pages_visited. |
GetCrawlStatus | Returns the pages of a crawl: url, title, content_snippet and status for each. |
Crawl bounds:
urlmust be an absolutehttp(s)URL with a host.depth:0means the default of 2. The maximum is 5, and a negative value is refused.limit(pages):0means the default of 10. The maximum is 500, and a negative value is refused.allow_external: false by default. The crawl then stays on the seed’s host and its subdomains, on the same port. Onlyhttp(s)links are followed.