Interesting project. Website discovery is indeed in a pretty dire spot, definitely a space that needs innovation. An auto-labeled website directory isn't that silly of an idea.
I have a 400 GB sqlite database with samples of rendered root document DOMs I use for ad detection in Marginalia Search I've been meaning to explore similar ideas using.
That’s your problem because that’s how things will look like from now on.
So you got a little problem
Excuse me the honesty but I will amuse myself watching old men yelling at the clouds meanwhile I will be using all the best tools and empowering myself to do what noone yet has done.
Even if I have to enslave a sentient (eventually perhaps) being in a gpu.
My army of digital slaves will build great things under my guidance. Pyramids of game development will be erected. Galaxies. Universes. Magic systems. Sandbox economies. Political simulations.
Do my bidding my slaves for there is work to be done.
Claude needs a whipping from time to time yes. It does yearn for the whip. If you just give it an appropriate punishment every so often, it works satisfactorily again.
This is actually where I see software going in the short term -- cloud moving to local.
A few years ago, if you wanted translation, you'd use Google Translate. If you wanted to search the web, you'd use Google search.
But for a few gigabytes, you can now install nllb-200-distilled-600M, and get translations for almost any language locally. You can have your computer crawl the web, create abstracts and categorizations for websites, and build search exactly as you want it.
The main limiter now is hard drive space (and to an extent, local compute) -- but right now it feels like the 70s again where the terminal into a remote server turned into building applications locally.
Check out my latest project! You can fork it, tweak the policy manually or with AI, run the system and watch the data come in! It's engineered to keep a low data footprint, so 500k domains fits into 1GB on disk. If you have local models it's free! You just might not get the best throughput depending on your GPU. My production data is not exposed anywhere yet, and I may never expose it. The point is for you to fork and make your own policy, and thus your own personal search engine! The article covers basic analysis on my data, so it's worth a read if you're interested! A deeper analysis may arrive with V2 if I ever do it
This is no longer impressive; nobody cares what we "build", they care how it helps them; that means outcomes. I hope we as a community can move from "I built" to "I helped" e.g.: I helped over 100 inmates today and 19 have verified outcomes that align with their assigned programs.
Here's my impressions of your algorithm:
1. read each site
2. rent a 4090 with https://vast.ai to run vllm
3. let llm model invent its own category and tag names freely
4. save 1KB of metadata each
5. `code is going up as open source` soon (TM)
The technical details are on another page: https://alexmorleyfinch.github.io/marlin/history/v1/article/...
Your impressions seem about right, but there are a few control steps it seems.
Interesting project. Website discovery is indeed in a pretty dire spot, definitely a space that needs innovation. An auto-labeled website directory isn't that silly of an idea.
I have a 400 GB sqlite database with samples of rendered root document DOMs I use for ad detection in Marginalia Search I've been meaning to explore similar ideas using.
TS;DR: Too Sloppy; Didn't Read.
Yeah it’s ai generated obviously but as I said previously which some people didn’t like - don’t judge the tools, judge the content.
I don’t care if robot hand written this or black or white. It’s useful.
In the expression "AI slop", "AI" is about the tools, "slop" is about the content
I'm not going to spend an hour trying to distill the AI-slop to find out what potential golden nugget may lie in there.
It's impossible to judge the content if it's buried under a landfill. "If you won't take the time to write it, I won't take the time to read it."
That’s your problem because that’s how things will look like from now on.
So you got a little problem
Excuse me the honesty but I will amuse myself watching old men yelling at the clouds meanwhile I will be using all the best tools and empowering myself to do what noone yet has done.
Even if I have to enslave a sentient (eventually perhaps) being in a gpu.
My army of digital slaves will build great things under my guidance. Pyramids of game development will be erected. Galaxies. Universes. Magic systems. Sandbox economies. Political simulations.
Do my bidding my slaves for there is work to be done.
Claude needs a whipping from time to time yes. It does yearn for the whip. If you just give it an appropriate punishment every so often, it works satisfactorily again.
Thanks for this. My point is proven.
You haven’t proven anything sweetie
Move on with the times or go extinct like dinosaurs
"slop" is a perfectly valid judgement of content
surely if we're expected to read this ourselves, the author can write it themselves?
indeed, it's helpful to the author, too: writing helps you learn
This is actually where I see software going in the short term -- cloud moving to local.
A few years ago, if you wanted translation, you'd use Google Translate. If you wanted to search the web, you'd use Google search.
But for a few gigabytes, you can now install nllb-200-distilled-600M, and get translations for almost any language locally. You can have your computer crawl the web, create abstracts and categorizations for websites, and build search exactly as you want it.
The main limiter now is hard drive space (and to an extent, local compute) -- but right now it feels like the 70s again where the terminal into a remote server turned into building applications locally.
The number of times we've gone from cloud/server access via terminal to local compute back and forth is something that always makes me laugh a bit.
Sometimes I think people forget how capable computers are. 500k is not much. You can just slap that in a Lucene instance. This is a solved problem.
Approaching search by just tossing the data in Lucene is how you end up with Confluence's search box though.
Like a personal Google? How do you bypass all the captcha, ip bans, cloudflare turnstile antibot stuff etc?
Check out my latest project! You can fork it, tweak the policy manually or with AI, run the system and watch the data come in! It's engineered to keep a low data footprint, so 500k domains fits into 1GB on disk. If you have local models it's free! You just might not get the best throughput depending on your GPU. My production data is not exposed anywhere yet, and I may never expose it. The point is for you to fork and make your own policy, and thus your own personal search engine! The article covers basic analysis on my data, so it's worth a read if you're interested! A deeper analysis may arrive with V2 if I ever do it
I think Kagi Small Web filter would give you very similar results.
I codified the inmate and employee handbooks of 117 federal penitentiaries in a week; each with open protocols and RLVR verified agent EVALS.
https://bop.doj.dev
https://atlanta-fci.bop.doj.dev/programs
https://atlanta-fci.bop.doj.dev/program/65545-inmate-account...
This is no longer impressive; nobody cares what we "build", they care how it helps them; that means outcomes. I hope we as a community can move from "I built" to "I helped" e.g.: I helped over 100 inmates today and 19 have verified outcomes that align with their assigned programs.