Class LuceneIndexManager
-
Nested Class Summary
Nested ClassesModifier and TypeClassDescriptionstatic final classEverything about this index that an operator can tune, in one object so adding a knob does not add a positional constructor argument. -
Field Summary
Fields -
Method Summary
Modifier and TypeMethodDescriptionbackup()Takes a consistent, hot backup of the index into a.ziparchive sibling to the index dir, namedl2-lucene-index.backup-{yyyyMMdd_HHmmss}.zip.voidclose()voidcommit()Commits pending changes now, instead of waiting for the scheduled commit.voidRemoves documents by url or tag.voidRemoves every document from the index.longRemoves all documents whose effective expiration timestamp is in the past.longOn-disk size of the index, summed over the files Lucene currently lists.intDocument count as a reader sees it - the number the management API and the console report.intDocument count, read from the live writer rather than through a searcher.static LuceneIndexManagergetInstance(String indexPath, double maxMergedSegmentMB, double deletesPctAllowed) Returns the one manager forindexPath, creating it on first use.static LuceneIndexManagergetInstance(String indexPath, LuceneIndexManager.Tuning tuning) Returns the one manager forindexPath, creating it withtuningon first use.getKeys()Returns all URL keys stored in the index.getKeys(int limit) At mostlimitkeys - the same scan asgetKeys(), stopped early.getResponse(String url) Gets response from index
-
Field Details
-
JSON_FIELD
- See Also:
-
BIN_FIELD
- See Also:
-
TAGS_FIELD
- See Also:
-
URL_FIELD
- See Also:
-
EXPIRE_DATE_FIELD
- See Also:
-
JSON_TYPE
public static final org.apache.lucene.document.FieldType JSON_TYPE
-
-
Method Details
-
getInstance
public static LuceneIndexManager getInstance(String indexPath, double maxMergedSegmentMB, double deletesPctAllowed) Returns the one manager forindexPath, creating it on first use.This is how production code must obtain a manager. Only one
IndexWritercan hold an index directory, so sharing a single manager per directory is what keeps a second, permanently broken writer from ever existing - seeINSTANCES.- Parameters:
indexPath- index directory (any form; normalised internally)maxMergedSegmentMB- cap on largest merged segment (MB); <=0 -> defaultdeletesPctAllowed- % deleted docs tolerated before reclaim; clamped to Lucene's [20,50]
-
getInstance
Returns the one manager forindexPath, creating it withtuningon first use.- Parameters:
indexPath- index directory (any form; normalised internally)tuning- merge, commit and refresh settings - seeLuceneIndexManager.Tuning
-
commit
public void commit()Commits pending changes now, instead of waiting for the scheduled commit.Writes on the request path deliberately do not commit (see
LuceneIndexManager.Tuning), which is the right trade for a cache: a lost write is re-fetched from origin. It is the wrong trade for an operator-initiated invalidation, where an entry re-appearing after a crash is a surprise rather than a warm-up - so the management-API paths that empty or purge the cache call this.Never throws: a wedged writer is discarded here exactly as it is anywhere else, and the caller's own operation has already succeeded in memory.
-
backup
Takes a consistent, hot backup of the index into a.ziparchive sibling to the index dir, namedl2-lucene-index.backup-{yyyyMMdd_HHmmss}.zip.The current writer is committed and a
SnapshotDeletionPolicysnapshot pins the commit's files so they survive concurrent merges while they are streamed into the archive. Each file is zipped on the fly (no intermediate copy on disk); entries are nested under a top-level folder that keeps the original index directory name (e.g.l2-lucene-index), so unzipping yields a directory identical to the live index - only the archive file name carries the.backup-{yyyyMMdd_HHmmss}suffix. The snapshot is always released afterwards so the pinned files can be reclaimed.- Returns:
- the name of the created backup archive
- Throws:
IOException- if the index writer is unavailable or the archive can't be written
-
close
public void close() -
delete
Removes documents by url or tag.- Parameters:
urlOrTag- a full URL (normalised to its path), a bare path, or an invalidation tag
-
deleteAll
public void deleteAll()Removes every document from the index. There used to be a domain-scoped flavour that matched a DOMAIN_FIELD term. Entries pushed in by index replication carry the SOURCE edge's domain, so it silently walked past them and they stayed served forever. With one site per node there is one flavour, and no DOMAIN_FIELD. -
deleteExpired
public long deleteExpired()Removes all documents whose effective expiration timestamp is in the past. "Cache forever" entries are indexed with Long.MAX_VALUE (see indexDoc()) so they are never selected by the expired range. Documents indexed before the EXPIRE_DATE_FIELD point existed carry no point value and are left untouched until they are re-cached.- Returns:
- number of expired documents removed from the index
-
getKeys
Returns all URL keys stored in the index.Uses a paginated
MatchAllDocsQuerywithIndexSearcher.searchAfter(ScoreDoc, Query, int)and loads only theURL_FIELDstored field per document, avoiding the cost of materialising large binary / JSON fields for every hit.- Returns:
- list of normalised URL keys (never
null)
-
getKeys
-
getResponse
Gets response from index- Parameters:
url- request url- Returns:
- WebResponse from index
-
getIndexSizeFast
public int getIndexSizeFast()Document count, read from the live writer rather than through a searcher.IndexWriter.getDocStats()is an in-memory read of state the writer already maintains - cheaper still thangetIndexSize()'s searcher acquire, and, being the writer's own view, it does not lag a refresh. That is what a Prometheus scrape arriving every few seconds forever wants. The two can differ by whatever has been written since the last searcher refresh; for a cache-size gauge that is noise.- Returns:
- the document count, or -1 when the writer is unavailable - never an exception, because the caller is a metrics gauge and a wedged index must not break a scrape
-
getIndexDiskBytes
public long getIndexDiskBytes()On-disk size of the index, summed over the files Lucene currently lists. This is the number that answers "will the cache volume fill up", which document count does not: entry sizes vary by orders of magnitude on a page cache. Deleted-but-not-merged documents are included, deliberately - they occupy the disk this is measuring.- Returns:
- total bytes, or -1 when the directory cannot be read
-
getIndexSize
public int getIndexSize()Document count as a reader sees it - the number the management API and the console report.Reads through the shared near-real-time searcher. It used to open its own
FSDirectoryandDirectoryReader, which read the last COMMIT - and since commits became periodic (seeLuceneIndexManager.Tuning) that would report a count up to one commit interval stale, and zero on a node that has not committed yet. The NRT view is both correct and far cheaper.- Returns:
- the document count, or -1 when L2 is unreadable
-