Give Google the boot by building your own search engine
If you're fed up with search results drowning out the bits of the web you actually care about, you could always build your own search index, as one developer did.
Nottingham, UK-based software dev Alex Morley-Finch built the open-source project dubbed Marlin for himself, cataloging around 560,000 homepages for roughly $10 in rented cloud GPU time and using less than a gigabyte of disk storage.
Morley-Finch said in a writeup of the project that he wanted a search engine of things he cared about, like “portfolios, zines, weird little art projects, one-person software,” and other cases of just “people doing stuff” that they share on the internet.
Something simple, with “a crawler that only ever looks at homepages, a small local language model that reads each one and writes a name, two or three sentences, a category, and a handful of tags,” he explained.
“No IP scanning, no Redis, no storing full page HTML, no recrawl scheduler, nothing multi-tenant.” Morley-Finch built Marlin around four processes: A fetcher that grabs domains, a worker that makes calls to a small, OpenAI-compatible language model, a steward that prevents bad pages from making their way into the index, and an API with a web UI where he can track the process and actually conduct searches of his index, complete with filters.
Of course, you can’t expect something like this to go perfectly on the first try.
“The first version worked within a couple hours.
Point it at a sample of domains, watch things get summarised, search for them.
Great,” he said in his writeup of the project.
“Sunday afternoon [he began working on Marlin on a Sunday], I looked at what had actually been catalogued and it was the wrong web.” Instead of indexing what he wanted it to, Marlin had just been grabbing a cross section of the internet, leading to more than 90 percent of the first attempt being “corporate sites and documentation.” Rather than blocking certain domains, Morley-Finch built a weighting system to push certain pages to the top of his crawl queue, and others as far down the list as possible.
Morley-Finch rented a cloud GPU to handle most of the heavy processing and stopped the crawl at around 560,000 pages after exhausting his prioritized categories and beginning to pull in "the raw internet." He spent around $10 on the cloud GPU.
The one overarching problem he had, and which he said may be a problem for anyone else who tries to replicate his project, is tagging and categorizing - left to a language model, those important elements got a bit messy.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.