Skip to content

perf(db): share collectors within predicate indexes - #424

Merged
jimezesinachi merged 6 commits into
mainfrom
perf/papaya-predicate-collector-sharing
Sep 21, 2026
Merged

jimezesinachi merged 6 commits into
mainfrom
perf/papaya-predicate-collector-sharing

Conversation

@jimezesinachi

Copy link
Copy Markdown
Collaborator

Summary

Reduce predicate-index memory usage by sharing one Papaya collector between the outer map and all value buckets belonging to the same predicate index.

Previously, every predicate-value bucket owned an independent collector. High-cardinality indexes could therefore create tens or hundreds of thousands of collectors, with the collector overhead dominating the memory used by the indexed values themselves.

The new per-index boundary keeps separate predicate indexes isolated while removing the collector-per-bucket cost.

Implementation

  • Introduce SharedPredicateIndex, which owns:
    • the outer MetadataValue -> HashSet<StoreKeyId> map;
    • one shared Seize collector;
    • inner bucket sets constructed from that collector.
  • Preserve the existing predicate lookup, insertion, and deletion behavior.
  • Keep the previous independent-collector implementation behind bench-experiments as the benchmark control.
  • Add fallible constructors for Papaya maps and sets using a shared collector.
  • Update DB size expectations for Papaya 0.2.5’s smaller map representation.

Persistence Compatibility

The collector is runtime state and must not become part of persisted snapshots.

SharedPredicateIndex therefore serializes exactly its inner map, preserving the previous serialized shape. During deserialization it:

  1. creates a new collector for the predicate index;
  2. reconstructs the outer map with that collector;
  3. reconstructs every value bucket with the same collector;
  4. restores the serialized store-key memberships.

Coverage verifies this behavior through:

  • direct predicate-index serialization round trips;
  • legacy JSON persistence migration;
  • standalone utils::Persistence recovery;
  • in-memory Raft snapshot recovery;
  • durable RocksDB-backed Raft snapshot recovery;
  • indexed reads and writes after restoration.

Papaya Safety Fix

Benchmarking shared collectors exposed an unsafe Papaya drop path: dropping one map called Collector::reclaim_all() even when another map sharing that collector still had an active guard.

The issue and reproducer are documented in:

The proposed upstream fix is:

Until that fix is merged and released, this PR pins Papaya to the exact reviewed commit:

1d914b94d846da6185380b814d09758624fd9f7e

Performance

Criterion compares the old independent-collector representation against the new per-index representation using the production insertion implementations.

Representative sequential results:

Workload Control Candidate Change
Construct empty index 478 ns 842 ns 76.2% slower
First unique 10k 14.239 ms 10.468 ms 26.5% faster
First unique 100k 243.30 ms 128.11 ms 47.3% faster
Low-cardinality 10k 1.482 ms 1.490 ms effectively neutral

The fixed construction cost increases by approximately 364 ns per predicate index. High-cardinality sequential ingestion improves substantially because it no longer creates an independent collector for every bucket.

Forced parallel high-cardinality insertion remains worse for the candidate:

Workload Control Candidate Change
Unique 10k parallel 59.209 ms 62.351 ms 5.3% slower
Unique 100k parallel 257.44 ms 478.96 ms 86.0% slower

The benchmark forces these paths for analysis. Production’s current 150k parallel threshold processes the measured 10k and 100k workloads sequentially.

Memory

Memory measurements run each control/candidate case in fresh processes and report the median of three runs.

Populated stores

Scenario Collectors C / P Allocated MiB C / P Saving
Empty 4 / 4 0.142 / 0.142 0.0%
1 index, low-cardinality 10k 15 / 5 1.579 / 1.420 10.1%
1 index, unique 10k 10,005 / 5 189.789 / 23.853 87.4%
4 indexes, unique 10k 40,008 / 8 755.925 / 91.992 87.8%
1 index, unique 100k 100,005 / 5 1,891.478 / 231.329 87.8%

After deleting store keys

Scenario Allocated MiB C / P Saving
1 index, low-cardinality 10k 1.256 / 1.095 12.8%
1 index, unique 10k 199.366 / 23.412 88.3%
4 indexes, unique 10k 795.905 / 90.912 88.6%
1 index, unique 100k 1,989.454 / 226.739 88.6%

The report also exposes immediate post-drop retention in the candidate’s deferred Papaya table retirement. For example, the unique 100k case retains approximately 225.6 MiB immediately after dropping the store. This is visible in the report for follow-up alongside the upstream reclamation fix.

Benchmark Reporting

Extend the opt-in benchmark-experiments workflow to support two report types:

  • Criterion control/candidate timing comparisons;
  • isolated control/candidate memory reports.

Changed memory targets under benches/memory/ generate Markdown that is appended to the existing update-in-place PR benchmark comment.

Validation

  • Targeted Clippy with -D warnings
  • Criterion and memory benchmark targets compile
  • Full isolated memory matrix completes
  • Serialization and restoration coverage added for standalone and clustered persistence

Signed-off-by: Jim Ezesinachi <ezesinachijim@gmail.com>
@jimezesinachi jimezesinachi added the benchmark-experiments Trigger opt-in Criterion benchmark comparisons label Sep 16, 2026
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Test Results

402 tests   402 ✅  11m 41s ⏱️
 42 suites    0 💤
  4 files      0 ❌

Results for commit b31f915.

♻️ This comment has been updated with latest results.

Signed-off-by: Jim Ezesinachi <ezesinachijim@gmail.com>
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Benchmark Results

group                                                        main                                   pr
-----                                                        ----                                   --
predicate_query_with_index/size_100                          1.00      2.6±0.01µs        ? ?/sec    1.00      2.6±0.00µs        ? ?/sec
predicate_query_with_index/size_1000                         1.00     25.3±0.03µs        ? ?/sec    1.02     25.9±0.06µs        ? ?/sec
predicate_query_with_index/size_10000                        1.01    251.8±0.71µs        ? ?/sec    1.00    248.2±1.04µs        ? ?/sec
predicate_query_with_index/size_100000                       1.00      4.0±0.02ms        ? ?/sec    2.39      9.5±0.58ms        ? ?/sec
predicate_query_without_index/size_100                       1.02      5.0±0.01µs        ? ?/sec    1.00      5.0±0.01µs        ? ?/sec
predicate_query_without_index/size_1000                      1.00     46.3±0.07µs        ? ?/sec    1.00     46.4±0.02µs        ? ?/sec
predicate_query_without_index/size_10000                     1.02    821.4±4.32µs        ? ?/sec    1.00    805.0±4.08µs        ? ?/sec
predicate_query_without_index/size_100000                    1.00     35.6±0.60ms        ? ?/sec    1.09     38.7±0.46ms        ? ?/sec
store_batch_insertion_without_predicates/size_100            1.01    157.7±0.84µs        ? ?/sec    1.00    156.5±2.03µs        ? ?/sec
store_batch_insertion_without_predicates/size_1000           1.00    947.8±5.66µs        ? ?/sec    1.00    951.1±3.31µs        ? ?/sec
store_batch_insertion_without_predicates/size_10000          1.00     10.6±0.07ms        ? ?/sec    1.02     10.8±0.10ms        ? ?/sec
store_batch_insertion_without_predicates/size_100000         1.00    107.6±1.44ms        ? ?/sec    1.02    109.5±1.96ms        ? ?/sec
store_retrieval_linear_cosine_similarity/size_100            1.02     17.1±0.04µs        ? ?/sec    1.00     16.7±0.04µs        ? ?/sec
store_retrieval_linear_cosine_similarity/size_1000           1.01    102.7±0.50µs        ? ?/sec    1.00    101.9±0.81µs        ? ?/sec
store_retrieval_linear_cosine_similarity/size_10000          1.00  1452.6±29.12µs        ? ?/sec    1.00  1455.7±28.87µs        ? ?/sec
store_retrieval_linear_cosine_similarity/size_100000         1.00     21.6±0.37ms        ? ?/sec    1.07     23.2±0.47ms        ? ?/sec
store_retrieval_linear_dot_product/size_100                  1.01     15.2±0.03µs        ? ?/sec    1.00     15.0±0.04µs        ? ?/sec
store_retrieval_linear_dot_product/size_1000                 1.01     97.1±0.37µs        ? ?/sec    1.00     95.9±0.43µs        ? ?/sec
store_retrieval_linear_dot_product/size_10000                1.00  1233.4±10.66µs        ? ?/sec    1.02  1253.0±15.73µs        ? ?/sec
store_retrieval_linear_dot_product/size_100000               1.00     20.0±0.33ms        ? ?/sec    1.02     20.5±0.13ms        ? ?/sec
store_retrieval_linear_euclidean_distance/size_100           1.02     15.7±0.01µs        ? ?/sec    1.00     15.4±0.05µs        ? ?/sec
store_retrieval_linear_euclidean_distance/size_1000          1.00     98.4±0.27µs        ? ?/sec    1.00     98.3±0.19µs        ? ?/sec
store_retrieval_linear_euclidean_distance/size_10000         1.00  1296.7±21.55µs        ? ?/sec    1.06  1373.7±37.04µs        ? ?/sec
store_retrieval_linear_euclidean_distance/size_100000        1.00     21.2±0.19ms        ? ?/sec    1.04     22.0±0.20ms        ? ?/sec
store_retrieval_no_condition/size_100                        1.00     16.8±0.01µs        ? ?/sec    1.00     16.9±0.03µs        ? ?/sec
store_retrieval_no_condition/size_1000                       1.02    103.3±0.80µs        ? ?/sec    1.00    101.6±0.43µs        ? ?/sec
store_retrieval_no_condition/size_10000                      1.10  1522.2±30.90µs        ? ?/sec    1.00  1379.4±19.74µs        ? ?/sec
store_retrieval_no_condition/size_100000                     1.00     22.1±0.19ms        ? ?/sec    1.00     22.2±0.15ms        ? ?/sec
store_retrieval_non_linear_hnsw/size_100                     1.00    124.7±0.04µs        ? ?/sec    1.00    124.2±0.36µs        ? ?/sec
store_retrieval_non_linear_hnsw/size_1000                    1.00    261.5±0.37µs        ? ?/sec    1.01    263.0±0.83µs        ? ?/sec
store_retrieval_non_linear_hnsw/size_10000                   1.00    417.6±7.13µs        ? ?/sec    1.06    441.3±3.71µs        ? ?/sec
store_retrieval_non_linear_hnsw/size_100000                  1.00    458.3±2.73µs        ? ?/sec    1.07    489.3±3.19µs        ? ?/sec
store_retrieval_non_linear_kdtree/size_100                   1.00    170.5±0.22µs        ? ?/sec    1.00    170.4±0.35µs        ? ?/sec
store_retrieval_non_linear_kdtree/size_1000                  1.00    961.7±2.61µs        ? ?/sec    1.00    964.6±7.29µs        ? ?/sec
store_retrieval_non_linear_kdtree/size_10000                 1.00     11.5±0.36ms        ? ?/sec    1.07     12.4±0.35ms        ? ?/sec
store_retrieval_non_linear_kdtree/size_100000                1.07    140.4±0.97ms        ? ?/sec    1.00    131.1±0.83ms        ? ?/sec
store_sequential_insertion_without_predicates/size_100       1.00    217.9±0.18µs        ? ?/sec    1.01    219.4±0.39µs        ? ?/sec
store_sequential_insertion_without_predicates/size_1000      1.01      2.2±0.00ms        ? ?/sec    1.00      2.2±0.00ms        ? ?/sec
store_sequential_insertion_without_predicates/size_10000     1.00     21.8±0.04ms        ? ?/sec    1.01     22.0±0.03ms        ? ?/sec
store_sequential_insertion_without_predicates/size_100000    1.00    219.9±0.45ms        ? ?/sec    1.00    219.4±0.49ms        ? ?/sec

Signed-off-by: Jim Ezesinachi <ezesinachijim@gmail.com>
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Opt-in Benchmark Results

Criterion results from the benchmark targets changed by this PR.

db/papaya_collector_sharing

first-time benchmark: candidate vs control

group                                                                         pr-35609073992-db-papaya_collector_sharing/candidate/    pr-35609073992-db-papaya_collector_sharing/control/
-----                                                                         -----------------------------------------------------    ---------------------------------------------------
papaya_collector_sharing/construction/1                                       1.84  1598.4±56.63ns        ? ?/sec                      1.00   869.8±20.65ns        ? ?/sec
papaya_collector_sharing/ingestion/existing_low_cardinality_10k_parallel      1.40  1498.7±80.44µs  6.4 MElem/sec                      1.00  1069.4±18.99µs  8.9 MElem/sec
papaya_collector_sharing/ingestion/existing_low_cardinality_10k_sequential    1.09  1591.1±50.13µs  6.0 MElem/sec                      1.00   1466.3±4.45µs  6.5 MElem/sec
papaya_collector_sharing/ingestion/first_low_cardinality_10k_parallel         1.00   977.1±10.32µs  9.8 MElem/sec                      1.08   1056.9±6.58µs  9.0 MElem/sec
papaya_collector_sharing/ingestion/first_low_cardinality_10k_sequential       1.00   1401.5±4.64µs  6.8 MElem/sec                      1.02   1435.7±2.59µs  6.6 MElem/sec
papaya_collector_sharing/ingestion/first_unique_10k_parallel                  1.00      6.8±0.27ms 1441.3 KElem/sec                    1.35      9.2±0.24ms 1065.8 KElem/sec
papaya_collector_sharing/ingestion/first_unique_10k_sequential                1.00     12.0±0.04ms 815.9 KElem/sec                     2.21     26.5±0.52ms 369.1 KElem/sec

db/papaya_collector_memory

Median of 3 isolated processes per case. Values are control / candidate.

Scenario Phase Collectors Allocated MiB Saving Active MiB Resident MiB
empty populated 4 / 4 0.107 / 0.107 0.0% 0.133 / 0.133 0.133 / 0.133
empty after deleting store keys 4 / 4 0.107 / 0.107 0.0% 0.133 / 0.133 0.133 / 0.133
empty after dropping store 0 / 0 0.107 / 0.107 n/a 0.133 / 0.133 0.133 / 0.133
one_index_low_cardinality_10k populated 15 / 5 1.296 / 1.132 12.7% 1.371 / 1.207 1.387 / 1.219
one_index_low_cardinality_10k after deleting store keys 15 / 5 0.974 / 0.812 16.6% 1.277 / 1.109 1.711 / 1.547
one_index_low_cardinality_10k after dropping store 0 / 0 0.720 / 0.658 n/a 1.020 / 0.953 1.711 / 1.547
one_index_unique_10k populated 10005 / 5 174.819 / 8.848 94.9% 174.910 / 8.910 177.832 / 9.086
one_index_unique_10k after deleting store keys 10005 / 5 184.630 / 8.416 95.4% 184.855 / 8.688 188.098 / 9.414
one_index_unique_10k after dropping store 0 / 0 0.402 / 8.261 n/a 3.316 / 8.531 188.098 / 9.414
four_indexes_unique_10k populated 40008 / 8 697.239 / 33.271 95.2% 697.387 / 33.406 711.059 / 34.094
four_indexes_unique_10k after deleting store keys 40008 / 8 737.232 / 32.233 95.6% 737.449 / 32.797 752.375 / 34.426
four_indexes_unique_10k after dropping store 0 / 0 0.138 / 32.089 n/a 5.090 / 32.648 751.785 / 34.426
one_index_unique_100k populated 100005 / 5 1744.860 / 84.719 95.1% 1744.992 / 84.852 1778.199 / 86.590
one_index_unique_100k after deleting store keys 100005 / 5 1842.885 / 80.193 95.6% 1843.078 / 80.855 1880.246 / 89.848
one_index_unique_100k after dropping store 0 / 0 0.105 / 79.054 n/a 6.336 / 79.715 1878.449 / 89.848

@deven96
deven96 self-requested a review September 18, 2026 11:52

@deven96 deven96 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Approved but pending removal of the bootstrap code in the core engine

Signed-off-by: Jim Ezesinachi <ezesinachijim@gmail.com>
Signed-off-by: Jim Ezesinachi <ezesinachijim@gmail.com>
Signed-off-by: Jim Ezesinachi <ezesinachijim@gmail.com>
@jimezesinachi jimezesinachi changed the title Share collectors within predicate indexes perf(db): share collectors within predicate indexes Sep 21, 2026
@jimezesinachi
jimezesinachi merged commit 75602bd into main Sep 21, 2026
10 checks passed
@jimezesinachi
jimezesinachi deleted the perf/papaya-predicate-collector-sharing branch September 21, 2026 17:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark-experiments Trigger opt-in Criterion benchmark comparisons

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants