Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
-
Updated
Aug 16, 2026 - Rust
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Give your Hermes agent the web as real sources, never a made-up answer — multi-provider search and extraction with an optional local, key-free DonSeTch option.
Use LLMs to robustly extract web data
Fully automated and hands-free, accurately extracting and understanding web content — powered by machine learning agents.
Windows desktop app for collecting, reviewing, ranking, and exporting Google Scholar search results.
Low-Cost Cross-Domain Web Structured Information Extraction using specialized LoRA adapters.
Replayable Browser Agent
基于Scala Akka的分布式主题网络爬虫
A swipeable shortlist of live San Francisco apartment listings, powered by Context.dev.
Self-hosted web scraping and Markdown extraction for AI agents
Hermes plugin to improve web search with SearXNG for higher-quality search results and Crawl4AI for LLM-optimized webpage extraction.
Automatic extraction of the information on local event from a webpage with Machine Learning
Give your agent the web as real sources, never made-up answers — an MCP server with 15 search and 9 extract providers plus an optional local, key-free DonSeTch integration.
Free local web search/extraction router for AI agents. Go CLI + MCP, BYOK/free-first routing, keyless DDGS/Scrapling fallback, setup writers and client guides.
A powerful and lightweight web scraping library with LLM extraction capabilities. This library combines web scraping with AI-powered content extraction using either OpenAI or OpenRouter APIs.
Predicting product recommendation score using the data available on the website of the client
Fast local Tavily-compatible web search and extraction backend for Hermes
Standalone Crawl4AI web extraction and bounded crawling plugin for Hermes Agent
Deterministic runtime cognition infrastructure for humans and AI agents — the same input yields the same SHA-256 across Python, JavaScript, Dart, Java and Kotlin.
Programming assignments for Web Information Extraction and Retrieval, FRI UL, 2021. PA1: standalone webcrawler of .gov.si web sites, PA2: approaches of the structured web data extraction, PA3: Data processing and indexing and Data retrieval.
Add a description, image, and links to the web-extraction topic page so that developers can more easily learn about it.
To associate your repository with the web-extraction topic, visit your repo's landing page and select "manage topics."