The two Loghound configsets are deliberately small. This page lists what they do not define and why, and how they commit and cache.
The stock configset loads extraction, clustering, language detection, templating and data-import libraries. Every one is an unauthenticated code path behind a request handler, and the extraction one in particular has a long history of remote code execution triggered by a hostile document. Loghound needs none of them, and not loading a library is a stronger control than not defining its handler.
No streaming, SQL, export, replication, extraction, debug-dump, terms, term-vector, analysis, spelling, suggest, browse or elevation handlers. Each removal is annotated in the configset with what it would have handed an attacker. The short version: the streaming and SQL handlers execute a server-side language with network and database sources; the export handler dumps the entire core in one response; the replication handler serves the raw index files; the terms handler enumerates the term dictionary, which on this schema means every address and every URL in the index without a query.
Remote streaming and stream bodies are written out as disabled explicitly even though they are already the defaults, so that copying a stock configuration over this one is a visible change rather than a silent regression — and the Solr client refuses to send any streaming parameter independently. Two locks on the same door.
The one catch-all that remains
A dynamic field that swallows anything not named in the schema and indexes none of it. The trade-off, stated plainly: a schema and code that drift apart degrade quietly — a value from a future version vanishes instead of failing the request, which is right for a running ingest daemon, because the alternative is one bad field killing an entire batch and stalling the tailer. The cost is that a typo in a field name is silent, and so is a live index left behind by an upgrade. If you are developing new fields, comment that line out and let Solr reject them loudly until you are done — and in production run bin/loghound-schema after every upgrade, which is the supported way to catch the same problem.
Both indexes use a five-second soft commit for visibility and a sixty-second hard commit that flushes segments to disk and truncates the transaction log without opening a new searcher. That second half is the important one: it gives durability with no cache flush and no warming stall, and visibility stays entirely the soft commit’s job. Getting it backwards is the classic way to make a Solr node spend all its time warming.
The two indexes differ on caches because their workloads are mirror images: hits is write-heavy, read-rare and append-only, so nothing is autowarmed there — with a five-second soft commit a searcher lives five seconds, and warming it costs more than the queries it would serve. Sessions is read-heavy, write-light and facet-dominated, with writes arriving once a minute and the dashboard reusing the same handful of filters, so warming pays for itself.
Both use the higher-compression codec rather than the faster default, because log lines are extremely repetitive: it typically reaches four to six times on the raw line where the default reaches two to three. The cost is processor time on stored-field retrieval, and stored fields are retrieved only in the drill-down, one screen at a time.
Reference
- Identifiers & labels
- Signal and reason codes
- Verdicts and bot classes
- Networks, referrers, crawlers
- The Solr schema
- Index size
- What the configsets leave out