Skip to content

[test] Isolate XQSuite test files, run CoreTests and XQuery3Tests in parallel - #6796

Open
duncdrum wants to merge 27 commits into
developfrom
test/xqsuite-isolation
Open

duncdrum wants to merge 27 commits into
developfrom
test/xqsuite-isolation

Conversation

@duncdrum

@duncdrum duncdrum commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Follow-up to the JUnit 5 migration, which is merged (#6790, #6791). Part of #6037.

Summary

#6791 added parallel XQSuite files (@XQSuite(parallel = true)) but almost no suite could use it: test files share one database and collide on collection names and global state. This PR makes the files independent, enables parallel runs where that is safe, fixes the product bugs and reporting gaps the work turned up, and records what still blocks the rest.

"Isolated" means 6-12 parallel runs at 2, 4 and 8 files at a time are green. Parallelism is only the tool used to find the collisions.

Needs a closer look (changes under src/main): items 2 (XSLT catalog), 7, 8, 10, 11 (Lucene), 12 (three new public system: functions) and 15 (EXPath registry). Items 5 and 6 change the test reports, including test names. Items 2, 5 and 6 do not depend on the isolation work and can move into their own PRs.

What changed

Reports

  1. [test] Publish the run time of each XQSuite file; one line per suite names the slowest.
  2. [test] Failure reports show the expected and actual value (built with Jupiter's AssertionFailureBuilder, cut at 4000 characters). Before, surefire dropped them.
  3. [test] Test names lead with their file (ft-query-field.xqm: issue5431-...) and carry the source line so an IDE can jump to the test. Reported names of XQSuite tests change, so CI history and flaky-test tracking restart.

Test engine (the first two commits; also open as #6823 and #6824, they drop out of this PR when one of those merges first)
19. [test] The engine's own fixture suites (the nested hang and failing-file classes of its tests) are skipped unless the test that sets them up asks for them (@XQSuite(fixture = true), configuration parameter exist.xqsuite.fixtures). A -Dtest wildcard that also matches nested classes, or an IDE package scan, ran them as ordinary suites and reported two failures for each of the hang fixtures after 5 minutes.
20. [bugfix] The wait for the threads of abandoned files is bounded by one grace period in all. It passed max(0, remaining) to Thread.join, and join(0) waits forever, so with two or more files stuck in a lock the run never ended (observed on #6775: XMLDBTests at 8 files at a time, still hung after almost two hours). Test: SuiteRunAwaitTerminationTest, three threads that ignore interruption; times out with the old formula.

Isolation
3. [test] Unique collection names per file, one commit per suite (CoreTests, XQuery3Tests, LogicalOpTests, SystemTests, OptimizerTests, RangeTests, FileTests, XmlDiffTests). Each file declares its collection name once, as a variable of the file; the rules for parallel suites are in the Javadoc of XQSuite#parallel().
12. [feature] Query-scoped trace functions system:enable-query-tracing, system:query-trace, system:clear-query-trace, used by the %test:stats handling. The database-wide trace made stats tests see other files' calls. The old functions are unchanged. Touches Profiler, QueryTrace, SystemModule, xqsuite.xql. Tests: QueryTraceTest (two overlapping queries) and two XQSuite guards for a stats test that throws.
17. [test] RangeTests queries are scoped to their own collection (part of the RangeTests commit): the six context-less range:field-eq calls of range.xql and the 17 range:index-keys-for-field calls of conditions.xql and prefixed-conditions.xql (without a context that function reads every document of the database). range.xql and conditions.xql also create a bystander collection with the same index and other data, so an unscoped query fails every time. On a loaded Windows CI runner one call found the copy of the same data stored by optimizer.xql (update-delete failed once; green in 30+ local runs). With the bystander present and without the scoping, 3 tests of range.xql and 5 of conditions.xql fail.
18. [test] FunDeepEqualPerformanceTest: the absolute timing caps become relative checks (fastest of 5 runs after 2 warm-ups). The 1500 ms root-mismatch test failed on earlier heads of this PR; it timed the construction of two in-memory trees (1019 to 1053 ms on three successful CI runs, longer than the equal-trees query) and could not detect a lost short-circuit. Now: stored equal documents at most twice a plain pass over them (about 0.29 locally), root mismatch on stored documents at most a tenth of the full comparison (1.3 ms against 155 ms), in-memory equal trees at most 5 times the construction alone. Checked by mutation: with the streaming path disabled the stored guard fails (3171 ms against 521 ms), with the root short-circuit removed the mismatch test fails (110.9 ms against 160.3 ms); the old caps passed both. The ratios on the CI runners are not measured yet.
4, 13, 14, 16, 21. [test] Parallel files enabled for CoreTests, XQuery3Tests (12 of 12 runs green each), RangeTests (24 of 24), OptimizerTests (12 of 12), ExpathRepoTests (12 of 12) and XMLDBTests (12 of 12 at 2, 4 and 8 files at a time, 4 reps each, 45 tests; it stayed serial because of the collection copy deadlock fixed in #6775, and without that fix the same setup deadlocked in 2 of 3 runs, so the suite also guards the lock order). NumericOpTests was already on.

Product fixes found on the way
2. [bugfix] XSLT imports: try the system catalog before the database resolver. The resolver fetched absolute http URLs with no connect timeout first; transform/catalog.xql waited 150 s on a non-routable address. XQuery3Tests 184 s -> 8 s (thread dump at 53 s).
7. [bugfix] Lucene ft:binary-field: read the value from the reader the search ran on, not a second snapshot (and compare with the segment size, not the live count). Was the flaky issue5431-sorted-by-binary-field (about 1 run in 5 at 8 files at a time, 0 of 40 after). Test: LuceneBinaryFieldLookupTest.
8. [bugfix] Lucene index scan (util:index-keys): fresh doc values iterators per term; they only move forward. Failed with an AssertionError on segments where only some documents have a node id. Test: LuceneIndexScanSparseDocValuesTest.
9. [test] Lucene tests no longer depend on score order or other files' documents; index-keys-by-qname tests use their own collection.
10. [bugfix] Lucene: analyze a document with the analyzers of the thread that adds it (FieldAnalyzerWrapper). One shared per-field registry let two threads swap each other's analyzer. Test: FieldAnalyzerWrapperTest (7). The interleaving was never observed directly; the evidence is a file pair that failed 16 of 74 runs before and 0 of 20 after.
11. [bugfix] Lucene ft:facets: keep the taxonomy reader alive as long as the facets can be read (AlreadyClosedException after a concurrent write and refresh). Held by a Cleaner, so an old reader stays open until a garbage collection. Test: LuceneFacetsReaderLifecycleTest.
15. [bugfix] EXPath package registry: serialize installs and removals. pkg-java 2.1.1 locks its package map only around parts of an install; choosing the directory, adding the version and the in-place rewrite of packages.xml / packages.txt run outside the lock, so concurrent calls read a half-written file (Error transforming the file: .../packages.xml) or could overwrite each other. ExistRepository now has installPackage / removePackage under a write lock (Deployment, PackageService, repo:install, repo:remove use them); a package that is not a local file is copied to a temporary file first so a slow download does not hold the lock; listPackages() returns a copy under a read lock because the library's own list is a live view that throws ConcurrentModificationException. Test: ExistRepositoryConcurrencyTest (5): concurrent install/remove failed 2 of 2 before and passes after; forced listing-versus-install interleaving; a stalled download does not block other installs or listings (fails when the copy is disabled); a failed download keeps its NotFoundException (pins existing behaviour). Limits: the Packages objects in a copy are still the library's; two installs of the same package run one after the other and the second replaces the first. The cleaner fix is in expath/expath-pkg-java.

Results

  • Parallel and green: see items 4, 13, 14, 16, 21.
  • Re-probed 2026-10-06 at 2, 4 and 8 files at a time, 2 runs each (6 per suite; 12 for OptimizerTests and ExpathRepoTests): every suite except the two blocked ones below is green in every run, with the same test count each time.
    • Parallel (7): CoreTests (70 files), XQuery3Tests (101), NumericOpTests (12), RangeTests (12), OptimizerTests (2), ExpathRepoTests (3), XMLDBTests (7; probed 2026-10-10 after the rebase, 12 runs).
    • Serial, isolated, 2 or more files (16): exist-core DateTests, IndexingTests, LogicalOpTests, MapTests, NumberTests, SecurityManagerTests, SystemTests, UtilTests, ValidationTests, XQSuiteTests; extensions LuceneIndexingTests, VectorSearchTests, NgramTests, CompressionTests, FileTests, XmlDiffTests. They run 0.6-7 s each, so parallel adds flake risk for no gain.
    • Serial, one file (14): nothing runs alongside them, so their runs say nothing about collisions: ArrayTests, InspectTests, XIncludeTests (13-14 s), and in extensions RestXqTests and XQDocTests (all tests skipped), IndexingTests (integration), FtAttrContextTests, FtMatchTests, LuceneAnalyzersTests, ReindexSecondStoreTest, SortTests, CacheTests, HttpClientXqueryTests, ImageTests.
  • CI, ubuntu unit job, "Maven Unit Tests" step (push / PR run, measured before the reporting and Lucene commits): develop 1163 / 966 s; [refactor] Replace XSuite with a JUnit Platform engine, remove JUnit 4, move test support to the test-jar #6791 1139 / 917 s; this PR 884 / 630 s. Runner noise between runs of the same code is 200-250 s, so only the difference of about 300 s means anything; almost all of it is XQuery3Tests (297-310 s on develop, 18-32 s here). The longest classes stay Java tests (DeadlockIT, CollectionLocksTest), so no module-level speed-up is claimed.

Still serial

Suite Cause Status
LuceneTests Defects found by parallel runs (items 7, 8, 10, 11) and the %test:stats files (item 12) are fixed. Rare failures remain at 4 and 8 files at a time: issue3977-reindex-doc-updates-not-adds (expected 1 2, got 1 1), some issue3042 / issue3043 sort tests, and now and then a cascade in ft-facets.xqm where one document's data stays wrong for the rest of the file. #6740 and #6756 were tried on top: no effect. Open, root cause not found; tracked in #6800 (the XMLDBTests row of that issue is resolved by #6775 and item 21).

Not addressed here: #4455 (facet counts of consecutive ft:query calls on the same nodes add up; its test stays %test:pending). Related: #1752 (does not reproduce), #4538.

Test plan

  • Rebased onto develop 198dabc0f3 (includes [refactor] Migrate tests from JUnit 4 to JUnit 5 #6790, [refactor] Replace XSuite with a JUnit Platform engine, remove JUnit 4, move test support to the test-jar #6791, [bugfix] omit newline after doctype with indent=no #5370 and [bugfix] Compare the full key when inserting into QNamePool #6795); getDocTypeIndentNo is a Jupiter test. Reactor test-compile; QNameTest, QNamePoolTest, NamePoolTest, SerializationTest (47, 0 failures); XQuery3Tests (1090, 0 failures, 3 skipped); RestoreAppsTest (6), PackageServiceJsonTest (4), ExpathRepoTests (9)
  • On the final history: CoreTests 981 (3 of 3 runs at 8 files at a time), XQuery3Tests 1090, OptimizerTests 60, XQSuiteTestEngineTest 13, RangeTests 411, FileTests 54, ExpathRepoTests 9, ExistRepositoryConcurrencyTest 5 and the new Lucene tests, all green; after every cleanup step the tree was identical to the tested tree
  • Each new guard test fails without its fix and passes with it (the registry stalled-download test checked by disabling the copy)
  • Serial runs of every consumer of the query-scoped functions green: OptimizerTests 60, QuantifiedMatchOptimizerTest 12, XQSuiteTests 93, XQSuiteTestEngineTest 13, CoreTests 981, LuceneTests 679, RangeTests 411
  • Each suite: serial run unchanged, 6 to 12 parallel runs green (above)
  • Rebased 2026-10-10 onto develop c68693484d (includes [bugfix] Lock the destination Collection first when moving or copying #6775): reactor test-compile of all modules, then org.exist.test.xqsuite.* (all engine tests, 34), XMLDBTests 45, MoveCopyLockOrderTest 7, WebDavMoveCopyLockOrderTest 5, all green; XMLDBTests parallel 12 of 12
  • CI on the current head 8ade722d3a (not run yet; the previous head d32ab7b80b was green on all jobs, the Windows integration job of an earlier head had failed RangeTests update-delete, fixed by item 17)

🤖 Generated with Claude Code

@duncdrum
duncdrum added this pull request to stack #6792 October 4, 2026 18:36
@duncdrum duncdrum added this to v7.0.0 Oct 4, 2026
@duncdrum duncdrum added this to the eXist-7.0.0 milestone Oct 4, 2026
@duncdrum

duncdrum commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor Author

The timing failure on one of the unit runs is interesting. It looks like the kind of flake that surfaces once a bunch of test timings have changed. So in a way its expected, and will likely disappear on rerunning. Will dig into this more. I m at roughly 19/38 testuites that go through the whole lets isolate you dance.

@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from ce5ff8d to 283f51f Compare October 4, 2026 19:09
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

📊 XQTS result comparison

Comparison of this run against develop.

Metric develop this run Change
➖ Passed 28,868 (90.74%) 28,868 (90.74%) 0 (0.00 pp)
➖ Failures 1,546 1,546 0
➖ Errors 134 134 0
➖ Skipped 1,267 1,267 0
🧪 Total tests 31,815 31,815 0

Relative to develop: 0 newly passing, 0 newly failing, 0 new errors, 0 newly skipped — counting only tests recorded in both runs whose outcome changed.

Runtime: 460.5s (+66.35s vs develop).

@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from 283f51f to 71da293 Compare October 5, 2026 09:24
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from 71da293 to fc34ffb Compare October 5, 2026 12:15
@duncdrum duncdrum added bug issue confirmed as bug Lucene issue is related to Lucene or its integration enhancement new features, suggestions, etc. labels Oct 5, 2026
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch 2 times, most recently from 5f91a45 to 00d12cf Compare October 5, 2026 16:07
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from 00d12cf to c882322 Compare October 6, 2026 06:58
Base automatically changed from refactor/exist-core-test-scope to develop October 6, 2026 10:03
@line-o
line-o force-pushed the test/xqsuite-isolation branch from 846f55c to fd58fce Compare October 6, 2026 10:03
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from fd58fce to 38620d6 Compare October 6, 2026 10:07
@duncdrum
duncdrum marked this pull request as ready for review October 6, 2026 10:52
@duncdrum
duncdrum requested a review from a team as a code owner October 6, 2026 10:52
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from f7453b4 to d32ab7b Compare October 6, 2026 10:57
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from d32ab7b to 366ccfa Compare October 7, 2026 07:47
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from 366ccfa to 23a86df Compare October 8, 2026 11:45
@dizzzz
dizzzz requested a review from a team October 9, 2026 08:32
@duncdrum duncdrum mentioned this pull request Oct 9, 2026
1 task done
duncdrum and others added 27 commits October 10, 2026 11:32
…s for them

The fixture suites of the engine tests are nested static classes annotated
with XQSuite. A run that selects them by name (a -Dtest wildcard such as
XQSuiteSchedulingTest*, or the package or classpath scan of an IDE) ran them
as ordinary suites. The hanging fixture then took the default five minutes
to be given up on, and the run reported two failures for each of
SequentialHang and ParallelHang, as well as the deliberately failing files
of XQSuiteTestEngineTest.

XQSuite gets a fixture flag. The engine skips such a suite unless the
configuration parameter exist.xqsuite.fixtures is true, which the engine
tests set on their EngineTestKit runs. A guard test checks that a fixture
is skipped without the parameter and fails without the engine change.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…riod in all

A test file that hangs in a database lock is given up on, but its thread
does not end when it is interrupted. SuiteRun then waits for the abandoned
threads before the database is shut down, and computed the time left for
each one as max(0, deadline - now) for Thread.join. The first stuck thread
uses up the whole grace period, and join(0) does not mean "do not wait",
it means "wait forever", so with two or more stuck files the run never
ended. The same happens for any thread when the grace period is set to 0.

The wait is now a helper that skips the threads once the time is spent,
and returns whether all of them ended. The tests use threads that ignore
interruption, and fail (time out) with the previous formula.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Surefire has no per-file entry. Publish the time of each file as a report entry,
and print one line per suite when it is done, with the number of files, the time
in all and the slowest files. A line printed as each file ends would land in the
output of whichever test is running at that moment, which with files running side
by side is a test of another file.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
EXistURIResolver opened absolute http URLs without a connect timeout before the
system catalog was consulted. transform/catalog.xql imports from the non-routable
address 203.0.113.1 and expects the catalog to redirect the import, so every run
of XQuery3Tests waited for the OS connect timeout: 150 s, with a thread dump at
53 s showing the thread inside the resolver's connect. XQuery3Tests went from
184 s to 8 s.

The resolvers are now tried in this order: EXPath package resolver, system
catalog, database resolver, default resolver. A catalog entry still wins over a
live network fetch (see #350). A product fix that does not depend on the test
isolation work.

Also drops an unused import from XsltURIResolverHelper.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
Files shared /db/coll, /db/test and /db/axes-test and removed each other's data
when run in parallel. Each file now uses a collection named after itself (for
example test-guest, test-npt, test-axes-persistent-nodes). Setup, teardown and
every reference are renamed together; no paths are rewritten in the harness.
Serial runs are unchanged.

axes.xml now removes its collection in its tearDown again: the removal was
commented out (it still named the old collection), unlike in the other XSuite
files; nothing else refers to /db/axes-xml-test.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
12 of 12 parallel runs green at 2, 4 and 8 files at a time; the suite takes
5.1 s serial and 2.6-3.6 s parallel. The rules the files must follow are in the
Javadoc of XQSuite#parallel(), which this commit extends.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
Files shared /db/xq3-test and /db/xquery3, so two of them running at the same
time stored and removed each other's documents. Each file now uses a collection
named after itself (for example xq3-higher-order).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
12 of 12 parallel runs green at 2, 4 and 8 files at a time; the suite takes
8.1 s serial and 5.1 s parallel.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
and.xml and or.xml, the two module-load-path files of the system suite and the
two optimizer files shared a collection each, so two of them running at the same
time would store and remove each other's documents. Each file now uses a
collection named after itself (for example test-module-load-path,
optimizertest-expressions).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
optimizer.xql used range.xql's collection range-test-range, with the same index
configuration and data, for its tests of nested placeName fields, so the two
files running at the same time shared documents. It now creates its own copy,
range-optimizer-copy.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
The three sync files (sync, sync-serialize, syncmod) shared the collection that
the shared fixtures hard-coded, /db/file-module-test. Each file now declares its
own (/db/file-module-test-sync, -sync-serialize and -syncmod) and passes it to
helper:setup-db and helper:clear-db, which take the collection as a parameter;
the collection variables of the fixtures are removed.

Also removes helper:sync-with-options, a wrapper around file:sync that nothing
calls, and the two sync files' unused child-collection variables.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
diff.xqm and compare.xqm used the same collection (test-xmldiff-compare), so
running at the same time they removed each other's documents. diff.xqm now uses
test-xmldiff-diff.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
A failed XQSuite assertion reached the surefire report as one line, for
example "XQuery failure: ft-query-field.xqm:352 issue5431-sorted-by-binary-field".
The expected and the actual value were carried by the AssertionFailedError, but
surefire keeps only the message and the stack trace, so a wrong result could not
be told from an empty one without running the test again by hand.

* Build the failure with Jupiter's AssertionFailureBuilder, which puts
  "==> expected: <x> but was: <y>" into the message and still carries the values
  for an IDE comparison. The second "at (file:line)" line of the message is gone;
  the headline already names the file and line, and the stack trace keeps them.
* Cut an expected or actual value at 4000 characters, saying how long it was,
  so that one very large result cannot flood the report. The same limit applies
  to the values an IDE shows for comparison.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
Surefire reports all tests of a suite under the suite class and uses the display
name of a test as its name, so tests of different files were indistinguishable
in a report, and equal names in two files collided.

* The display name of a test now leads with its file, for example
  "ft-query-field.xqm: issue5431-sorted-by-binary-field". The report names of
  XQSuite tests change.
* Give each test the line its function is declared on, as the position of its
  file source. Before, an IDE opened the test file at the top; now it jumps to
  the test. The line comes from util:inspect-function in the discovery query.
  XML tests, and tests found by compiling the module, have no line and keep the
  file only.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
ft:binary-field returned an empty value for some nodes, which showed as a flaky
issue5431-sorted-by-binary-field in LuceneIndexingTests when its files ran side
by side. LuceneIndexWorker.getBinaryFieldByExistDocId had two defects:

* It found the Lucene document number with a search on the searcher, and then
  read the value through a second, separately acquired reader. The searcher and
  the reader refresh independently, and a number is only valid in the snapshot
  it was found in, so a commit by another thread in between could point the
  lookup at another document. The value is now read from the reader of the
  searcher that found the number.
* It compared the number within a segment with the number of live documents
  (numDocs) instead of all documents (maxDoc). A segment with deleted documents,
  which a reindex or the deletion of another document leaves, has more numbers
  than live documents, so the documents numbered above the count were not
  found. The segment is now found with ReaderUtil.subIndex and the number is
  used as it is.

The new test stores a document that is then deleted before one that is kept, so
that the kept one is numbered above the live documents of the segment, and reads
the binary field of every item of it. It fails on the unfixed code (item 012 and
the two after it have no value) and checks that its precondition holds, so it
cannot pass without testing the lookup.

With the fix, 40 runs of LuceneIndexingTests with its files running side by side
(8 at a time) had no failure; before, about 1 in 5 runs failed.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
index-keys-by-qname-all and index-keys-by-qname-s (and the pending attribute-type test)
called util:index-keys-by-qname without a context, which scans the terms of every
document in the database and compares them with the terms of this file's own
documents. Any other file that indexed a p element at the same time added terms,
so the result depended on which other files were running. They now run on the
documents of the file's own collection, like the test for a collection that does
not exist beside them.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
LuceneIndexWorker.doScanIndex (util:index-keys and util:index-scan) read the doc
id and the node id of each hit from doc values iterators that it fetched once per
segment. The hits of every term start again at the first document, but a doc
values iterator only moves forward: advanceExact needs a target at or above the
current one. A segment in which every document has the value tolerates going
back; a segment in which only some do, which happens once documents without a
node id (written by ft:index) share it with indexed nodes, does not. With
assertions on, as in tests, the scan failed with an AssertionError in
IndexedDISI.advanceExactWithinBlock; with assertions off it can return wrong
values. It showed in LuceneTests when files ran side by side
(index-scan-path-lucene, index-scan-qname-lucene and their -has-terms tests).

Each term now gets its own iterators.

The new test writes a Lucene document without a node id into the segment of three
indexed items, checks that the segment is in that state, and scans terms whose
hits start before the last hit of the term before. It fails on the unfixed code
with the AssertionError.

Also drop the FIXME that said the doc id iterator was null and should not be: both
places that write documents add the doc id doc values, and the stored field
fallback is kept for older indexes. This was not verified at run time.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
search-title-text and search-title-text-db returned the right two documents in
the other order when other files indexed documents at the same time: ft:search
orders its hits by score, and the score depends on the statistics of the whole
index (document frequencies and the average field length), which other files
change. The tests check which documents match, so the uris are sorted, in these
two and in the two other tests of the file that expect two uris.

search-title-text-db searches the whole database ("/db/") on purpose, so it now
keeps only the hits in its own collection and no longer fails on the documents
of other files.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
The index and its IndexWriter are shared by all collections, and the analyzer of
a field is looked up by the name of the field, which does not tell collections
apart: every collection that indexes p elements writes into the same field, with
the analyzer of its own index configuration. LuceneIndexWorker.write() put the
analyzer into one shared map and then called addDocument, but the writer analyzes
the document when addDocument gets to it, and addDocument can wait for a flush
that another thread started. If a thread that indexes into another collection
registered its own analyzer for the same field in the meantime, the document was
analyzed with that one: with a collection that does not fold diacritics next to
one that does, "Ruesselsheim" with an umlaut was indexed unfolded next to the
folded term, and a stemming English collection lost its stemming.

The index writer now uses FieldAnalyzerWrapper, which takes the analyzers a thread
sets for the duration of its own addDocument call (LuceneIndex.addDocument) before
it looks into the shared map. The shared map stays as the registry for everything
else, and is now a ConcurrentHashMap, because it is written by several threads.

Found with analyzers-diacritics.xql and analyzers-english.xql, which index p with
different analyzers and run side by side at 2 files at a time: the pair failed in
16 of 74 runs without the change and in 0 of 20 with it. The swap itself happens
inside addDocument and could not be observed, so FieldAnalyzerWrapperTest models
it: it registers an analyzer from another thread while a document is being added.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
The facets of a search (ft:facets) are computed from the taxonomy reader of the
searcher the query ran on, but that searcher is handed back to the manager as
soon as the query returns. The index is shared by all collections, so a write
by another thread followed by a refresh replaces the searcher, and the old
taxonomy reader was closed: reading the facets then failed with "this
TaxonomyReader is closed". Serial runs never refresh between the query and
ft:facets, so only parallel runs (and a server with concurrent writers) saw it.

The facets now hold a reference on their taxonomy reader, which a Cleaner
releases once the facets are unreachable. LuceneFacetsReaderLifecycleTest
reads the facets of a search after another document was stored and the index
refreshed; it fails with the error above without this change.

Also drops three unused imports from LuceneIndexWorker.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
system:enable-tracing, system:trace and system:clear-trace work on the
statistics of the whole database instance. While tracing is on, every query
records into them, and a query that reads or clears them reads or clears what
all the others recorded. A %test:stats test that runs next to another test file
therefore sees calls it did not make, or an empty trace after the other file
cleared it, and the test files of a suite cannot run in parallel.

The new functions system:enable-query-tracing, system:query-trace and
system:clear-query-trace use the statistics that every query keeps for itself
(imported modules and evaluated queries share them with the query that loaded
them), never touch the database-wide ones, and do not write the global
tracelog setting. The three existing functions are unchanged, so monitoring that
reads the database-wide trace is not affected.

The XQSuite harness uses the new functions for %test:stats, and switches tracing
off when a test throws: the call that normally does it is never reached then,
and tracing stayed on for the rest of the file's query. QueryTraceTest runs
two queries at the same time and checks that a query that never switched
tracing on records nothing while another has it on, that the database-wide
flag stays off, and that clearing one query's trace keeps another's. Two tests
in xqsuite-tests.xql check that nothing is recorded after a stats test that threw.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
The files of this suite used the database-wide performance trace for their
%test:stats tests, so parallel runs failed in the files that check the index
statistics. With the query-scoped trace functions that is gone: 24 of 24 runs
were green at 2, 4 and 8 files at a time (the suite takes about 4 s either way,
so the point is to keep it independent, not to save time).

Parallel runs also exposed queries that were not scoped to a file's own
collection. On the Windows CI job range.xql update-delete failed once, after more
than 30 green local runs: six calls of range:field-eq without a collection
context found the copy of the same data that optimizer.xql stores, because both
files index the field address-name and store the same names. The calls are now
collection($rt:COLLECTION)//range:field-eq(...). range.xql also creates a
bystander collection with the same index and documents, as another file would,
so a query that is not scoped fails here every time and not only when another
file happens to run at the same moment: without the scoping update-delete,
update-replace and update-value fail with the bystander present.

range:index-keys-for-field() called without a context reads every document of the
database. The 17 calls in conditions.xql and prefixed-conditions.xql were safe
only because each field name is defined in exactly one file; they are scoped as
well, and conditions.xql has a bystander with the same index and other key
values: with the calls unscoped 5 tests fail (index-ends-with, index-eq-numeric,
index-lt-numeric, index-lt-non-numeric, index-starts-with); scoped, the suite is
411 tests, 0 failures, 3 of 3 runs at 8 files at a time. The two tests of
index-keys.xql that call without a context on purpose, and the two wrapper
functions there, are left as they are, with a comment.

Also drops the unused variable $rt:DATA2 of range.xql.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
The suite has two files, expressions.xqm and positional.xqm, each with its
own collection since the unique-name commit. Its %test:stats tests used the
database-wide trace, which is gone with the query-scoped trace functions:
12 of 12 runs were green at 2, 4 and 8 files at a time (60 tests each; the
suite takes about 4 s either way, so the point is to keep it independent,
not to save time).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
All packages of a database share one registry: the package map of the wrapped
Repository and the files packages.xml and packages.txt. The library (EXPath
pkg-java 2.1.1) locks its package map only around parts of an install or removal;
choosing the directory of the package, adding the version to its Packages and the
rewrite of the registry files run outside that lock, and the files are rewritten
in place (read, transform, truncate, write). Two installs or removals at the same
time could read a half-written file ("Error transforming the file:
.../packages.xml"), and could overwrite each other's change and drop a package
from the registry. Any server with concurrent package installs could hit this; it
is what kept the two ExpathRepoTests files that install packages from running in
parallel.

ExistRepository now has installPackage (from an XarSource or a URI) and
removePackage (all versions or one) under a read-write lock, and the call sites
(Deployment, PackageService, the repo:install and repo:remove functions) use them
instead of the wrapped Repository. The whole call is the critical section, not
only the file write, because choosing the directory key and adding the version
are also unlocked in the library. A package that does not come from a local file
is first copied to a temporary file before the lock is taken, and installed from
the copy through a small source that still reports the original address, so a
slow download does not hold up other installs; the temporary file is deleted
afterwards. A failure to read the source is reported as the repository reports it
("Error unziping the package" with the IOException, which carries the
NotFoundException or HttpException of a download). installPackage(URI) delegates
to the XarSource overload and keeps the read-only storage check.

The library returns a live view of its package map from listPackages(), after
releasing its own lock, so a caller that does work per package (the module
lookups read files for each one) failed with a ConcurrentModificationException
when a package was installed or removed during the iteration.
ExistRepository.listPackages() returns a copy taken under the read lock, and the
callers that iterate (ExistRepository, Deployment, ClasspathHelper,
PackageService, repo:list, repo:get-resource, repo:resource-available) use it.
getParentRepo() documents why not to use the repository's own list or to change
the registry through it.

ExistRepositoryConcurrencyTest (embedded server, 5 tests): six threads installing
and removing packages of their own at the same moment (failed 2 of 2 on the
unlocked code with the production error, green in every run with the lock); the
registry on disk lists exactly the installed packages; a listing is not disturbed
by an install during the iteration (forced interleaving; fails with
ConcurrentModificationException on the repository's own list); a download that
sends its headers and stalls does not hold up another install or a listing (fails
after 30 s when the download is read under the lock, checked by disabling the
copy); a failed download reports a PackageException that carries the
NotFoundException (pins the existing error, does not fail without the change).

Not covered: the Packages objects in a copy are the library's own, so the
versions of one package can still change under a reader; two installs of the same
package run one after the other and the second replaces the first (that is what
force means).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
The two files that install and remove packages (deployment.xql and
resource-available.xql) failed together with "Error transforming the file:
.../packages.xml": 2 of 4 runs at 2 files at a time. The package registry
changes are serialized now (the two preceding commits), and 12 of 12 runs of
the suite were green at 2, 4 and 8 files at a time, 9 tests each, on the code
of this branch (the suite takes about 1 s either way, so the point is to
keep it independent, not to save time).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
FunDeepEqualPerformanceTest failed on earlier heads of this PR and could not
be trusted: its three timing tests compared one run against a fixed number of
milliseconds, and the 1500 ms cap of the root-mismatch test measured
something other than what it claimed. That test builds two 10,000-element
in-memory trees and compares them, and the comparison stops at the first
element name, so it times the construction. On three successful CI runs
(Ubuntu unit job) it took 1019, 1053 and 1025 ms, 68 percent of its cap, and
longer than the equal-trees query that does the whole comparison (557 to
792 ms); locally the mismatch (48 ms), the equal trees (49 ms) and the
construction alone (51 ms) are indistinguishable, so no threshold could have
detected a lost short-circuit. The stored-docs guard for GH-4050 ran 1140,
1589 and 2712 ms on CI against a cap of 5000 ms and about 155 ms locally.

Each test now compares two figures measured one after the other on the same
machine, each the fastest of 5 runs after 2 warm-ups (noise only makes a run
slower):
- stored equal documents: at most twice a plain pass over the same two
  documents (count of nodes and attributes); about 0.29 times locally;
- root mismatch: now on stored documents of the same size (a third document
  with another root name), where nothing has to be constructed, at most a
  tenth of the full comparison; 1.3 ms against 155 ms locally. The in-memory
  root mismatch keeps a correctness test without timing;
- in-memory equal trees: at most 5 times the same query without the
  comparison; about 1.05 times locally.

Validated by mutation, main code reverted afterwards:
- streaming fast path disabled (FunDeepEqual.tryStreamingCompare returns
  null): the stored guard fails, 3171 ms against 521 ms for the pass, where
  the old cap of 5000 ms would have passed;
- no short-circuit at a root name mismatch (both documents are read to the
  end before the mismatch is returned, result unchanged): the mismatch test
  fails, 110.9 ms against 160.3 ms, where the old cap of 1500 ms would have
  passed;
- unmutated, all 15 tests pass in two runs.
Not yet seen on CI, and not run under heavy machine load: the ratios on the CI
runners are unmeasured (the earlier scratch measurements under 4 busy processes
showed the same ratios for the same queries, but that was not these tests).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CLyS76KK3RFSauL2JEqqiY
The suite was left serial because xmldb:copy-collection deadlocked with
another file's test when the files ran side by side: the copy took the
read lock of the source Collection before the write lock of the target
(the parent of the source), and the other tests held intention locks on
/db. The lock order of copying and moving a Collection was fixed in #6775.

The seven files use collection names of their own, so they do not disturb
each other. Twelve runs on this branch (four reps each at parallelism 2, 4
and 8) passed 45 of 45 tests, none hung. Without the lock order fix the
same configuration deadlocked in two of three runs, so the parallel suite
also guards the lock order.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@duncdrum
duncdrum force-pushed the test/xqsuite-isolation branch from 23a86df to 8ade722 Compare October 10, 2026 10:40

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug issue confirmed as bug enhancement new features, suggestions, etc. Lucene issue is related to Lucene or its integration

Projects

Status: In review

Development

Successfully merging this pull request may close these issues.

2 participants