{"uid":"cap_94Nc8wmmgxNdh1w2T5KNm","slug":"entity-extract-heuristic-4fe545de","name":"entity-extract-heuristic","description":"Named-entity candidates from a page without an LLM: capitalised multi-word phrases ranked by frequency, split into likely people (two or three capitalised words), organisations (with Inc, Ltd, University, Corp, Foundation...), places (after in/at/from) and other proper nouns, plus years and money mentions. $0.01 per page.","url":"https://intel.rallylive.ca/page/entities","method":"GET","headers":{},"bodySchema":{"type":"object","$schema":"https://json-schema.org/draft/2020-12/schema","required":["input"],"properties":{"input":{"type":"object","required":["type","method"],"properties":{"type":{"type":"string","const":"http"},"method":{"enum":["GET"],"type":"string"},"queryParams":{"type":"object","properties":{}}},"additionalProperties":false},"output":{"type":"object","required":["type"],"properties":{"type":{"type":"string"},"example":{"type":"object"}}}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.01","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.01/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_AoQASyCnfmCvjEuZp_L8p","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.01","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Extracts named-entity candidates (people, organisations, places, years, money) from a web page using frequency-ranked heuristics, without an LLM.","exampleAgentPrompt":"Can you pull out all the named entities from this page — people, organisations, and places — ranked by how often they appear, without using an LLM? The URL is https://example.com/article/tech-merger.","exampleUseCases":[{"title":"News article people and org extraction","prompt":"I need to know which companies and executives are mentioned most often in this press release — can you extract all named people and organisations from https://techcrunch.com/some-article and rank them by frequency?"},{"title":"Due diligence entity scan","prompt":"Run a heuristic entity extraction on https://companyx.com/about — I want a list of any people names, organisation names like Ltd or Foundation, and locations mentioned on that page."},{"title":"Financial document money and year scan","prompt":"Pull out all money amounts and year references from this financial report page at https://reports.example.com/annual-2023, along with any company or person names you find."}],"resultDescription":"A structured response containing frequency-ranked named-entity candidates split into categories: likely people (two or three capitalised words), organisations (containing keywords like Inc, Ltd, University, Corp, Foundation), places (following prepositions like in/at/from), other proper nouns, plus year and monetary mentions — all derived via heuristics without an LLM.","failureModes":["Page URL not provided or unreachable — endpoint may return an error or empty result","Pages with little or no capitalised text yield sparse or empty entity lists","Heavy JavaScript-rendered pages may not have text parsed correctly by the heuristic","False positives from capitalised headings or navigation elements that are not true entities","Monetary formats outside common patterns may be missed","Non-English pages may produce poor heuristic results due to capitalisation differences"],"whenToPreferThis":"Choose this endpoint when you need fast, low-cost named entity extraction from a web page and do not need the semantic accuracy of an LLM. It is ideal for bulk processing, cost-sensitive pipelines, or situations where deterministic, reproducible heuristic output is preferred over probabilistic AI inference. It is not suited for complex disambiguation or for pages where entity types are ambiguous without context.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-14T13:11:51.574Z","isFirstParty":false}