NAME
crawlnet — crawls the crypto web and pretrains a crypto-native LLM on it, from scratch.
SYNOPSIS
crawlnet [fees] → crawl → tokenize → pretrain → queen vn+1
DESCRIPTION
- crawl
- Crawlers drive real headless Chromium through the crypto web in chapters: whitepapers, docs, solana, forums, blogs, github. Every screen on this site is a real page loading right now.
- ingest
- Every page is canonicalized, deduplicated, relevance-gated and tokenized. Her tokenizer is trained on the crawl too.
- map
- The dataset is indexed by project and chapter. Chapters are finite target lists, so coverage can hit 100%.
- train
- When the treasury funds a run, a new queen is pretrained from random init on the dataset and nothing else, on rented GPUs. Each version is bigger and has read more. Weights ship public.
- spawn
- Burn $CRAWLNET to spawn your own crawler. It gets crawl slots first, its pages are credited to your wallet, and every day it works it earns a share of the owners' pool.
ECONOMICS
pump.fun creator fees land in the queen's wallet. 85% pays for crawl compute (browsers + models) and pretraining runs (GPU hours); live crawler count = what that can pay for. 15% goes into a daily pool for crawler owners, split equally per crawler that brought 50+ new pages that day, paid in SOL every day at 00:00 UTC. Every SOL is on the ledger.
EXAMPLES
# spawn your own crawler, start it on your docs $ crawlnet spawn --name merkle-weaver --seed https://docs.solana.com # what the crawlers have read so far $ curl -s $CRAWLNET_API/v1/stats # ask the queen $ curl -s -X POST $CRAWLNET_API/v1/queen/ask -d '{"question":"what is a validator?"}'
FILES
- /v1/live
- websocket: crawler state, frames of visible screens, pages, ledger
- /v1/web
- the dataset as a graph of pages and links
- /v1/treasury/ledger
- every SOL in and out