Skip to content

Fix GPU and Java-heap OOM when building large vector indexes - #2476

Open
nvzm123 wants to merge 24 commits into
NVIDIA:mainfrom
nvzm123:zackm_cuvslucene-139
Open

nvzm123 wants to merge 24 commits into
NVIDIA:mainfrom
nvzm123:zackm_cuvslucene-139

Conversation

@nvzm123

@nvzm123 nvzm123 commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This ports NVIDIA/cuvs-lucene#173 into java/cuvs-lucene following the move into the cuVS monorepo.

The scope is limited to the accelerated-HNSW memory-pressure change:

  1. Float, binary-quantized, and scalar-quantized accelerated-HNSW construction uses native host-backed cuVS matrices instead of retaining a complete device-backed input matrix.
  2. All three writers include active primary host-input storage in codec-level ramBytesUsed() and emit allocation diagnostics through InfoStream. This observability does not update Lucene's cached IndexWriter counters or impose a native-memory limit.
  3. Ordinary FLOAT32 force merges stream Lucene's merged vectors directly into a native host matrix without creating an intermediate List<float[]>.
  4. Higher-layer construction can read selected rows from that native matrix without recreating the complete dataset on the Java heap.
  5. Empty and one-vector accelerated-HNSW fields are handled before allocating a non-trivial native build matrix, including ordinary FLOAT32 merge results.

The resulting codec still uses GPU CAGRA construction followed by Lucene CPU HNSW search.

Merge and ownership safety

The memory change includes the safety required by its new allocation and replay path:

  1. CuVSMatrix.Builder is closeable. The built-in builders release untransferred storage when closed, transfer ownership only after a successful build(), reject use after close or build, and allow repeated close() calls.
  2. A default no-op close() preserves compatibility for providers compiled against the earlier builder interface.
  3. Deletion-free merges use the merged-vector size reported by Lucene; merges with deletions count the live vectors returned by Lucene.
  4. The replay pass rejects both fewer and more vectors than the first pass observed and closes the builder on read, population, validation, or build failure.
  5. The existing public AcceleratedHNSWUtils.createMultiLayerHnswGraph(...) list-based descriptor remains available; the native-matrix overload used internally is package-private.
  6. Writer-owned native datasets close if CAGRA construction fails and transfer to the built index only after success. Temporary subset indexes and upper-layer adjacency matrices close after their data has been copied.
  7. Empty and singleton quantized fields avoid native matrix allocation, and native -1 padding is not serialized as an HNSW neighbor.

Scope boundaries

  • CuVS2510GPUSearchCodec remains device-backed and is otherwise unchanged by this PR.
  • Binary- and scalar-quantized force merges still materialize their existing heap lists; streaming those paths is deferred to a follow-up.
  • Upper-layer selection and remapping behavior is unchanged. The default remains one HNSW layer.
  • Broader native-resource cleanup beyond these owned build resources, GPU-search fallback policy, malformed-graph validation, and upper-layer correctness hardening are deferred to a follow-up.

Validation

Validation completed during development of this branch:

  • cuvs-java: 5 unit tests and 111 integration tests; 0 failures, 0 errors, and 1 inherited skip
  • cuvs-lucene: 348 tests; 0 failures, 0 errors, and 30 inherited skips
  • Three-level reopen-and-search coverage for float, binary-quantized, and scalar-quantized accelerated HNSW
  • Deleted, sparse, zero-live-vector, and one-live-vector force-merge coverage
  • Merge replay underflow, overflow, cleanup, and suppressed-failure coverage
  • Host-input accounting and diagnostics, dataset ownership transfer, cleanup failures, and singleton no-allocation/no-neighbor coverage
  • Spotless, generated-documentation idempotence, and git diff --check

Deep1B 1M, 96 dimensions, using fresh index directories:

  • Float, binary-quantized, and scalar-quantized builds each flushed four 250,000-vector segments and force-merged them into one three-level accelerated-HNSW segment.
  • For every format, validation reopened the retained index and checked checkIntegrity(), 1,000,000 documents and vectors, a complete unique source-ID domain, graph level membership and nesting, duplicate-free adjacency, finite search results, and ground-truth overlap.

These Deep1B runs are correctness sanity checks. Page-cache eviction was unavailable in the validation container, so the recorded timings are not performance evidence.

Retest Results

Retested the live PR-2476 head eb6c03e77 on Deep1B 100M at 96 dimensions. Both layouts used the historical cold-source configuration, forceMerge=0, and 1,000 measured queries.

Layout Revision Indexing Recall@1500 Mean search latency
1 -> 1 Previous 6c175502 595.929 s 94.945% 12.039 ms
1 -> 1 Current eb6c03e77 587.283 s 94.988% 12.115 ms
4 -> 4 Previous 6c175502 502.563 s 96.699% 25.380 ms
4 -> 4 Current eb6c03e77 504.464 s 96.710% 25.953 ms

Signed-off-by: Zack Meeks <zmeeks@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Aug 17, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@nvzm123
nvzm123 marked this pull request as ready for review August 20, 2026 04:03
@nvzm123
nvzm123 requested a review from a team as a code owner August 20, 2026 04:03
rapids-bot Bot pushed a commit that referenced this pull request Sep 10, 2026
This PR replaces [cuVS-Lucene #195](NVIDIA/cuvs-lucene#195) as cuVS-Lucene has been merged into cuVS.

This PR builds on the out-of-core host-streaming refactor from #2476 by @nvzm123. This PR subsumes #2476 to avoid stacking the PRs.

On top of #2476, this PR adds the following.

**Improvements**

- Native flat buffering: stream vectors directly into a native host matrix during indexing instead of buffering them as a heap List<float[]>, cutting peak host memory from ~2x to ~1x and eliminating the per-vector matrix-assembly copy. Opt-in; requires all input vectors to be indexed in the original order, i.e. not supported for sorted, merged, or filtered index segments; not currently supported for segments with quantized fields.
- Parallelized CAGRA-to-HNSW graph conversion: materializing the CAGRA adjacency into on-heap NeighborArrays was a serial per-node loop; parallelized under the existing `writerThreads` knob.
- Parallelized level-0 HNSW graph serialization: level-0 (all N nodes) is now delta/VInt-encoded in parallel across memory-bounded waves of threads, then concatenated to the `IndexOutput` in node order according to on-disk format in the serial path.

**New example: OptimizedCagraHnswBuildExample**

A reference pattern for building a large accelerated HNSW index with all ingest- and build-side optimizations, showcasing:

- Streaming, prefetched, bounded-memory ingestion: open the source file once, read it front-to-back in large sequential chunks; hold at most two chunks in memory; fill the next chunk on a background thread while the ingest thread drains the current one, hiding disk read behind indexing; unpack into a caller-reused float[] allocation (no per-vector allocation, safe because Lucene copies the value eagerly inside `addDocument`).
- Native flat buffering sized exactly to each segment's vector count via `withNumInputVectors`, avoiding the heap-buffered assembly copy.
- Overlapping multi-segment build:
    - Default: K single-segment passes appended to one directory; peak host memory is one slice (N/K).
    - Overlapped: a bounded pool (`PIPELINE_DEPTH`) builds segments concurrently into their own directories with the GPU commit serialized on a semaphore (ingest overlaps a prior segment's GPU commit), then combines the finished per-segment indexes by hardlinking their files into the final directory via `HardlinkCopyDirectoryWrapper` + `addIndexes` (no bulk copy of vector data). Peak host memory is up to `PIPELINE_DEPTH * N / K`.
- An `enableRMMAsyncMemory()` call (with a note that it must not be used with CPU-only codecs) to opt-in to RMM-managed memory resources, plus exposing the primary tuning knobs (withMaxConn, withBeamWidth, withCuvsDistanceType, withWriterThreads).

Authors:
  - James Xia (https://github.com/jamxia155)
  - https://github.com/nvzm123
  - Igor Motov (https://github.com/imotov)

Approvers:
  - James Lamb (https://github.com/jameslamb)
  - Igor Motov (https://github.com/imotov)

URL: #2481
@cjnolet cjnolet added improvement Improves an existing functionality non-breaking Introduces a non-breaking change labels Sep 11, 2026
Acquire device streams before allocation, preserve ownership across builder cleanup, and keep CAGRA persistence on the original device matrix. Expand lifecycle, fallback, merge, and graph-integrity regression coverage.
Keep host-backed CAGRA-to-HNSW inputs, exact live-vector merge sizing, trivial merge handling, and the compatible upper-layer bridge. Restore the GPU-search codec to the target-branch device-input behavior and defer broader lifecycle, graph-integrity, and quantized-merge hardening to a follow-up.
@imotov

imotov commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

/ok to test 8faae50

@imotov imotov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That looks much better. I think we are really close. Left a few minor comments.

import org.apache.lucene.util.InfoStream;

/** Observes the real codec writer, not IndexWriter's cached buffering counters. Requires cuVS. */
public class TestAcceleratedHNSWHostInputMemory extends LuceneTestCase {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you ensure that this an all other tests that require GPU to run as skipped on CPU-only machines?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done! The GPU-dependent tests added or updated here now skip when cuVS isn’t available. If cuVS reports support, they still require an accelerated writer, so silent CPU fallback can’t produce a false pass.

/** Counts the live vectors that the merge iterator will actually yield. */
private static int countMergedVectors(FieldInfo fieldInfo, MergeState mergeState)
throws IOException {
FloatVectorValues mergedVectors =

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: I think these iterations can be skipped and we can got with total of original sizes if there were no deletes or updates (all liveDocs are empty).

@nvzm123 nvzm123 Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great suggestion, thanks — this now uses mergedVectors.size() when every liveDocs entry is null and only iterates when a source has deletions.

writer.commit();

try (DirectoryReader sourceReader = DirectoryReader.open(writer)) {
assertEquals("the test requires three source segments", 3, sourceReader.leaves().size());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That fails for me with seed 5A68E1972548005F:36610B2E43B74A02

To reproduce run

LD_LIBRARY_PATH=$PWD/../../cpp/build/c:$PWD/../../cuvs/cpp/build${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH} \
mvn -B test \
  -Dtest=TestAcceleratedHNSWDeletedDocuments \
  -Dtests.method=testForceMergeCountsOnlyLiveSparseVectors \
  -Dtests.seed=5A68E1972548005F:36610B2E43B74A02

For this check we should probably disable auto flush like we do in other tests so it doesn't interfere with segments.

@nvzm123 nvzm123 Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed! The fixture now disables document-count auto-flush and uses a 256 MB RAM buffer with NoMergePolicy, so explicit commits control the source segments. The reported seed passed, along with 10 repeated class iterations..

@imotov

imotov commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

/ok to test 3b0f875

@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

/ok to test 3b0f875

@imotov, there was an error processing your request: E2

See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/

@imotov

imotov commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

/ok to test 83a5f97

@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

/ok to test 83a5f97

@imotov, there was an error processing your request: E2

See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/

@imotov

imotov commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

/ok to test 83a5f97

@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

/ok to test 83a5f97

@imotov, there was an error processing your request: E2

See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/

@imotov

imotov commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

/ok to test 83a5f97

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
@nvzm123
nvzm123 requested a review from imotov October 2, 2026 16:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

improvement Improves an existing functionality non-breaking Introduces a non-breaking change

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

4 participants