Skip to content

Web and crawler

FACE reads the web in two ways:

  • a web connection reads one configured page as rows of text
  • the crawler service scrapes a page, or runs a bounded crawl from a seed URL

Web connection

Type spellingswebtelemetry (web_telemetry), webpage (the “Webpage” card), web, url, website
Settings (properties)url (also endpoint, target_url, page_url, website, host)
Schemeshttp and https only
Internal addressesRefused by default. A page named by a URL is normally public. Set allow_internal_network=true to read an intranet page.
  • Execute fetches exactly one page and follows no links:
    • 30-second timeout, body capped at 8 MiB.
    • At most 5 redirects, on the same host only.
    • Only text and HTML are accepted. A document or image is refused, with a pointer to the file path.
    • FACE extracts the readable text and splits it into chunks of up to 2,048 characters, with at most 400 chunks.
  • Rows have the columns url, final_url, fetched_at, http_status, content_type, title, chunk_index, chunk_count, chunk_chars, content, truncated, source and transport. truncated is true when the body or the chunk count hit its cap.
  • Test fetches the page, reading at most 64 KiB, within 15 s, and applies the same address and redirect rules.
  • Explore returns supported=false. The connector reads one page by design.

The request’s User-Agent is Mozilla/5.0 (compatible; RuninkFACE/1.0) unless the operator sets SCRAPER_USER_AGENT.

Crawler service

CrawlerService exposes:

RPCBehaviour
ScrapeReads one url and returns its title, main content and metadata. render_js is reserved.
CrawlStarts from url and visits linked pages, returning a crawl_id, a status and pages_visited.
GetCrawlStatusReturns the pages of a crawl: url, title, content_snippet and status for each.

Crawl bounds:

  • url must be an absolute http(s) URL with a host.
  • depth: 0 means the default of 2. The maximum is 5, and a negative value is refused.
  • limit (pages): 0 means the default of 10. The maximum is 500, and a negative value is refused.
  • allow_external: false by default. The crawl then stays on the seed’s host and its subdomains, on the same port. Only http(s) links are followed.