Skip to content
Crawler, metasearch & CRM

Crawler, metasearch & CRM

Web collection and CRM access. CrawlerService scrapes and crawls public web pages. MetasearchService runs the sovereign metasearch engine and exposes the prompt guardrail’s rule sets for inspection. HubspotService reads deals and contacts from HubSpot. All three run inside the FACE deployment: no third-party search or AI API is involved.

Summary

ServiceRPCKindPurpose
CrawlerServiceScrapeUnaryFetch and extract one page
CrawlerServiceCrawlUnaryCrawl from a seed URL and record the pages
CrawlerServiceGetCrawlStatusUnaryRead a finished crawl’s pages
MetasearchServiceSearchTermsUnaryRanked, content-extracted web search
MetasearchServiceEvaluateBiasUnaryCheck text against a guardrail rule set without blocking
MetasearchServiceListBiasRulesUnaryList a guardrail rule set
HubspotServiceConnectUnaryVerify a HubSpot key
HubspotServiceGetDealsUnaryRead deals from HubSpot
HubspotServiceGetContactsUnaryRead contacts from HubSpot

CrawlerService

Full name semantics.v1.CrawlerService.

Scrape

rpc Scrape(ScrapeRequest) returns (ScrapeResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: INVALID_ARGUMENT if url is empty. A page that cannot be fetched, or no free browser slot for render_js, returns success: false with error_message instead of an error.

Fetches one page and extracts its title and main content. With render_js: true, the page is rendered in the server’s headless browser first, which waits for a browser slot. Otherwise it is fetched as plain HTML.

Request: ScrapeRequest

FieldTypeDescription
urlstringRequired. The page to fetch.
render_jsboolRender the page in a headless browser before extracting. The proto comment calls this reserved for future use, but the server implements it.

Response: ScrapeResponse

FieldTypeDescription
successboolWhether the page was fetched and extracted.
urlstringThe URL that was fetched.
titlestringPage title.
contentstringMain detected content, as Markdown or plain text.
metadatamap<string, string>Not populated by the server.
error_messagestringWhy the scrape failed, when success is false.

Crawl

rpc Crawl(CrawlRequest) returns (CrawlResponse);
  • Kind: Unary. The call returns when the crawl has finished.
  • Auth: Bearer session.
  • Errors: INVALID_ARGUMENT if url is empty or not an absolute http(s) URL with a host, or if depth or limit is negative. If the call’s deadline passes or the call is cancelled mid-crawl, the matching status is returned (DEADLINE_EXCEEDED or CANCELLED).

Crawls breadth-first from the seed. Only http and https links are followed. The crawl stops when limit pages have been attempted. The result is kept under the returned crawl_id in the serving process’s memory, for GetCrawlStatus.

depth counts link hops from the seed: the seed is at depth 0, and links are followed from any page whose depth is below depth. For example, depth: 1 fetches the seed and the pages it links to.

Request: CrawlRequest

FieldTypeDescription
urlstringRequired. Seed URL. It must be an absolute http(s) URL with a host.
depthint32Maximum link hops from the seed. 0 means the default, 2. Capped at 5.
limitint32Maximum pages to attempt. 0 means the default, 10. Capped at 500.
allow_externalboolFollow links to other hosts. When false (the default), the crawl stays on the seed’s host or its subdomains, on the same port. This is a host rule, not a registrable-domain rule: with seed www.example.com, a link to example.com counts as external.

Response: CrawlResponse

FieldTypeDescription
crawl_idstringId for GetCrawlStatus.
statusstringCompleted, or Failed when at least one page was attempted and none was fetched. The call is synchronous, so the server never returns Started.
pages_visitedint32Pages fetched successfully.

GetCrawlStatus

rpc GetCrawlStatus(GetCrawlStatusRequest) returns (GetCrawlStatusResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: None beyond authentication. An unknown crawl_id returns status: "Unknown" and no pages.

Returns the pages recorded for a crawl. Crawl results are held in memory by the process that ran the crawl. They do not survive a restart and are not shared between replicas.

Request: GetCrawlStatusRequest

FieldTypeDescription
crawl_idstringThe id returned by Crawl.

Response: GetCrawlStatusResponse

FieldTypeDescription
statusstringCompleted, Failed or Unknown.
pagesrepeated CrawledPageEvery page attempted, in visit order.

MetasearchService

Full name semantics.v1.MetasearchService.

This service returns only numbers that the search engine or the guardrail computed. It reports no invented score, ratio or percentage.

SearchTerms

rpc SearchTerms(SearchTermsRequest) returns (SearchTermsResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: INVALID_ARGUMENT if query is empty or only whitespace. UNAVAILABLE if the search fails.

Runs a live search through the sovereign metasearch engine. The engine drives a headless browser, fetches and extracts candidate pages, then ranks them. This is an expensive call, not a typeahead: bind it to an explicit user action.

Request: SearchTermsRequest

FieldTypeDescription
querystringRequired. The search query. Leading and trailing whitespace is removed.
max_resultsint32Ranked results to return. 0 or negative means 5. Capped at 20.

Response: SearchTermsResponse

FieldTypeDescription
querystringThe query as searched, echoed so responses can be matched to requests.
resultsrepeated SearchTermResultRanked hits. There may be fewer than requested.
requested_resultsint32How many results the engine was asked for, after the default and cap were applied.

EvaluateBias

rpc EvaluateBias(EvaluateBiasRequest) returns (EvaluateBiasResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: INVALID_ARGUMENT if agent_name is not a known rule set (the message lists the known names), or if text is empty or only whitespace.

Checks text against an agent’s guardrail rule set without blocking anything, and reports every rule that fires. The enforcing guardrail stops at the first rule that fires. It uses the same matching, so this result matches what enforcement would do.

Request: EvaluateBiasRequest

FieldTypeDescription
textstringRequired. The text to evaluate, such as a hypothesis, scenario name or search term.
agent_namestringThe rule set to use. See Rule sets. Empty means HypothesisService. An unknown name is refused rather than silently falling back.

Response: EvaluateBiasResponse

FieldTypeDescription
agent_namestringThe rule set actually used, after the default was applied.
rules_evaluatedint32How many rules the text was checked against. Read hits against this number.
hitsrepeated BiasRuleEvery rule that fired, in evaluation order. Empty means the text passes.
blockedbooltrue when the enforcing guardrail would refuse this text. Currently this is exactly “any hit”. Read this field rather than computing it from hits.

ListBiasRules

rpc ListBiasRules(ListBiasRulesRequest) returns (ListBiasRulesResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: INVALID_ARGUMENT if agent_name is not a known rule set.

Returns the full rule set an agent’s prompts are checked against, so you can see what the gate checks without tripping it.

Request: ListBiasRulesRequest

FieldTypeDescription
agent_namestringSame names, default and refusal as EvaluateBiasRequest.agent_name.

Response: ListBiasRulesResponse

FieldTypeDescription
agent_namestringThe rule set returned.
rulesrepeated BiasRuleThe full rule set, baseline rules first. This is also the order in which a first-match refusal picks its rule.

Rule sets

agent_name accepts exactly these values:

AnalysisService, ClaimsService, ComplianceService, CopService, DocumentationService, FetchService, FinanceService, FulfilmentService, HypothesisService, PostureService, PredictiveMaintenanceService, ReverseLogisticsService, RulesService, SelfHealingService, SurveillanceService, TelemetryService, TwinsService, UnderwritingService, VisionService, VoiceService.

HubspotService

Full name semantics.v1.HubspotService.

HubSpot access uses a key that the operator configures on the deployment. A key sent to Connect is only verified: it is never stored or used for later calls.

Connect

rpc Connect(HubspotServiceConnectRequest) returns (HubspotServiceConnectResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: INTERNAL if the verification request cannot be built. Every other outcome is reported in success and message.

Verifies api_key by making one read-only request to HubSpot. The key is not stored. For GetDeals and GetContacts to work, the operator must configure a HubSpot key on the deployment.

Request: HubspotServiceConnectRequest

FieldTypeDescription
api_keystringHubSpot API key or private-app access token. Write-only: never returned, logged or stored.

Response: HubspotServiceConnectResponse

FieldTypeDescription
successbooltrue if HubSpot accepted the key.
messagestringOutcome, including HubSpot’s HTTP status on rejection.

GetDeals

rpc GetDeals(GetDealsRequest) returns (GetDealsResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: FAILED_PRECONDITION if no HubSpot key is configured on the deployment. A HubSpot transport, HTTP or decoding failure is returned without a specific status code (clients see UNKNOWN).

Returns one page of deals from HubSpot’s default page size, with the properties below.

Request: GetDealsRequest

FieldTypeDescription
pipeline_idstringNot applied by the server. Deals from all pipelines are returned.
limitint32Not applied by the server.

Response: GetDealsResponse

FieldTypeDescription
dealsrepeated HubspotDealThe deals.

GetContacts

rpc GetContacts(GetContactsRequest) returns (GetContactsResponse);
  • Kind: Unary.
  • Auth: Bearer session.
  • Errors: Same as GetDeals.

Returns one page of contacts from HubSpot’s default page size.

Request: GetContactsRequest

FieldTypeDescription
limitint32Not applied by the server.

Response: GetContactsResponse

FieldTypeDescription
contactsrepeated HubspotContactThe contacts.

Messages

CrawledPage

FieldTypeDescription
urlstringPage URL.
titlestringPage title. Empty on error.
content_snippetstringThe first 500 bytes of extracted content, followed by ... when truncated.
statusstringSuccess or Error.

SearchTermResult

FieldTypeDescription
titlestringPage title.
urlstringPage URL.
contentstringClean readable text extracted from the page, truncated by the engine. Not a summary: no model has processed it.
scoredoubleThe engine’s relevance score for this hit. Comparable within one response only. Do not display it as a confidence or a percentage.
image_urlstringA representative image for the page, when the engine found one.

BiasRule

Used both for a rule that fired and for a rule in a rule-set listing.

FieldTypeDescription
rule_idstringRule id.
descriptionstringWhat the rule checks.
severitystringCRITICAL, HIGH, MEDIUM or LOW, verbatim from the rule.
categorystringOWASP, OpenBias or DomainScope, verbatim from the rule.
compliancerepeated stringCompliance standards the rule is cited under, for example ISO42001_AIManagerSystem or SOC2_AuditAndPrivacy.
is_domain_scopebooltrue for a must-match (domain scope) rule. A hit on such a rule means the text fell outside the agent’s domain, not that it contained something forbidden. Display the two kinds differently.

HubspotDeal

FieldTypeDescription
idstringHubSpot deal id.
dealnamestringDeal name.
amountstringAmount, as HubSpot returns it.
pipelinestringPipeline id.
dealstagestringDeal stage id.
closedatestringClose date, as HubSpot returns it.

HubspotContact

FieldTypeDescription
idstringHubSpot contact id.
firstnamestringFirst name.
lastnamestringLast name.
emailstringEmail address.