YaCy improved-search 日本語

An experimental YaCy fork: better ranking, signed results, peers behind NAT

A fork of YaCy that fixes weaknesses measured on the public network, adds a trust layer so that forged and spam results can be filtered, and lets peers behind a NAT take part through a relay. Every claim on this page is measured in closed peer-to-peer networks that anyone can rebuild with docker compose.

Status: experiment. This is not a release of the YaCy project and is not affiliated with it. It deliberately breaks compatibility with the public YaCy network (peer ids, seeds, CJK word hashes). It is shared to show what the changes do, with data, so that the ideas can be discussed and, where they fit, proposed upstream in small pieces.

Why

Measured on the public YaCy network (freeworld) with 16 queries, only 11% of the top 10 results contained all query terms. Looking into the causes:

What changed

Ranking

Stricter minimum match, term coverage weighting after per-peer normalization, a weight against thin pages, and searching all peers of small networks.

CJK

Chinese, Japanese and Korean text is indexed and searched as overlapping bigrams, in Solr and in the word index.

Trust

Ed25519 peer keys, signed seeds, coordinator-signed trust lists with declared tags, and an author signature on every document.

NAT traversal

A small go-libp2p sidecar reserves a slot on a circuit relay, so peers behind a NAT answer searches.

Search quality

SymptomCauseChange
Pages that match one term (keyword stuffing) rank on topSolr mm=1 for multi-term queriesMinimum match 2<-1 5<80%: two terms must both match, 3–5 terms may miss one (search.ranking.solr.mm, .mm.cjk)
A peer's best partial match ranks like other peers' exact matchesPer-peer score normalizationMultiply the normalized score by (terms found / query terms)² (search.ranking.coverage.exponent)
Japanese / Chinese not found in the word indexNo word segmentation for CJKOverlapping bigrams in the word index and the query; CJKWidthFilter + CJKBigramFilter in the Solr schema
Thin pages with the whole query in the title (tag lists) rank on topqf weighs title^15 and h1^11 against text^1Results with fewer than 100 words are weighted by words / 100 (search.ranking.thin.words); CJK word counts fixed (they counted spaces)
A network of new peers never searches other peers' word indexesDHT search needs peers older than 3 daysConfigurable (remotesearch.dht.minage, default 3)
Small networks send remote Solr queries to nobody, or skip DHT targetsTarget count formula gives 0; DHT targets were excluded from SolrNetworks of up to 32 peers ask every connected peer, DHT targets included

Trust layer

NAT traversal

A sidecar process (Go, go-libp2p) runs next to YaCy with the same key. Behind a NAT it reserves a slot on a circuit relay v2 and announces the circuit address in the signed seed (Reach=relay). Other peers open a local tunnel port to it and use plain HTTP, so YaCy's existing clients work unchanged. By default such peers only answer searches. They do not store DHT data unless they opt in.

See the trust and NAT design for the details.

Results

Two experiments in yacy-lab. Both run closed networks in docker compose, with a deterministic corpus and query set.

Search quality: upstream vs fork, 3 peers each

Each peer crawls one site. Queries go to peer 1 with resource=global, and most relevant pages are on the other peers. The corpus contains two kinds of decoys: pages stuffed with one query term, and thin "tag archive" pages with the whole query in the title. Mean over 11 queries (6 English, 4 Japanese, 1 Chinese), 2 runs with the same result except where noted.

Cluster / pathR-precision ↑Recall@10 ↑Decoys in top R ↓All terms in top 10 ↑
upstream, default0.520.960.480.42
fork, default0.84–0.881.000.12–0.160.75
upstream, word index only0.020.020.000.09
fork, word index only0.930.950.070.77

Trust and NAT: 6 fork peers, a relay and a NAT

Three trusted peers, a trusted peer that declares ads, a signed but untrusted peer that crawls spam and plants documents with a borrowed signature, and a peer behind a MASQUERADE router. All 26 checks pass, among them:

Try it

The demo starts both networks (3 upstream peers, and the fork's trust and NAT setup) on one machine and puts a search page in front of them. You need Docker with about 7 GB of memory.

git clone https://github.com/pad01g/yacy_search_server.git yacy
git clone https://github.com/pad01g/yacy-lab.git
cd yacy
git checkout baseline        && docker build -t yacy-lab/upstream:baseline -f docker/Dockerfile .
git checkout improved-search && docker build -t yacy-lab/fork:latest -f docker/Dockerfile .
docker build -t yacy-lab/sidecar:latest sidecar/
cd ../yacy-lab
docker compose -f compose.demo.yaml -p yacydemo up -d
# open http://localhost:8800 (setup takes about 10 minutes and shows its progress)
The demo page: the same query sent to the upstream network (left) and the fork (right). Upstream lists spam pages first; the fork lists verified relevant pages first and, in open mode, the spam below them labelled unverified.
"bitcoin lightning channel" in open mode. Left: upstream shows spam first. Right: the fork shows verified results first and the spam, labelled unverified, below.

The page also lets you change trust: choose which coordinators the searching peer trusts (a second coordinator lists only the spam peer, so trusting it makes the spam "verified"), edit and re-sign the trust list, hand the new version to one peer and watch it spread, and revoke or restore the operator's delegation.

The experiments themselves: docker compose -p yacylab up -d && docker compose -p yacylab run --rm runner (search quality) and docker compose -f compose.trust.yaml -p yacytrust up -d && docker compose -f compose.trust.yaml -p yacytrust run --rm runner (trust and NAT). See the lab README.

For AI agents

An agent can run its own peer and use it as a search tool, without a search API or API key:

Limits