Skip to content

[pull] main from llvm:main - #5820

Open
pull[bot] wants to merge 4443 commits into
Ericsson:mainfrom
llvm:main
Open

pull[bot] wants to merge 4443 commits into
Ericsson:mainfrom
llvm:main

Conversation

@pull

@pull pull Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

@pull pull Bot locked and limited conversation to collaborators Sep 16, 2026
@pull pull Bot added the ⤵️ pull label Sep 16, 2026
xiongzile and others added 28 commits October 9, 2026 12:05
Fixes #225361.

This patch fixes two issues related to bit-fields:

- **Big-endian CodeGen:** Clang incorrectly places the value bits of
oversized bit-fields after the padding bits, contrary to the Itanium C++
ABI (§2.4). Fix the layout so that value bits precede padding bits.
- **`__builtin_clear_padding` (LE and BE):** Correct the occupied-bit
calculation for bit-fields, including `bool` and `_BitInt`, by using
`min(declared width, type size)`. This preserves bits that should not be
treated as padding.
…beled DO (#230308)

When an OpenACC loop or combined construct is associated with a labeled DO
loop, AccNonBlockDoConstruct turns a terminating END DO statement into a
labeled CONTINUE statement, which silently dropped any construct name on the
END DO. The DO statement of an unnamed labeled DO loop has no construct
name, so a name on its END DO is not allowed (C1135), and it is already
diagnosed when the loop is not associated with a directive:

```fortran
!$acc parallel loop
do 10 i = 1, n
  a(i) = 0
10 end do foo
```

Report "Unexpected DO construct name" at the name before it is dropped.
Also add tests for labeled END DO forms that are accepted: a branch to the
END DO from inside the loop, END DO followed by an end directive, the kernels
and serial loop constructs, and a named labeled DO loop whose END DO
specifies the same name.

Assisted-by: AI
Lanes converted from mul-by-power-of-2 are emitted as shl by the
exponent; check the node shift amounts, not the scalar operands.

Fixes #230392

Reviewers: 

Pull Request: #230464
#223709)

The majority of the changes are just a case of ensuring the chain is
routed correctly and the matching STRICT passthrough node is used.

NOTE: At present full strict-fp support has a minimum requirement of
+sve2+bf16, otherwise we lack the necessary cast instructions. Of these
+bf16 is fundamental whereas +sve2 is only required for double->bfloat.
`ConditionalEvaluationFinder` in `CIRGenCleanup` skipped implicit code
because that's the default for `RecursiveASTVisitor`. This lead to
default arguments and default member initializers being skipped.

This patch enables the traversal of implicit code with the exception of
the implicit call to `await_resume()` in `co_await` and `co_yield`
expressions. This requires cleanup scopes for await full-expressions
which don't exist yet.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Use VectorInstrContext to return more accurate scalarization overhead
costs when inserts/extracts can be folded into ld1/st1 and CPUs where
ld1/st1 are fast (same perf as regular loads).

Depends on #175982

PR: #177201
…rs (#223540)

`DecodeThumb2BCCInstruction` omits predicate operands for Thumb barriers
when the data-barrier feature is disabled, causing llvm-objdump
--arch-name=thumb to abort.

Add the missing operands using the current IT state, with disassembler
and llvm-objdump regression tests.

Fixes #193877
Unnamed IR values are hard to follow in pass-by-pass output. Their
printed
numbers can shift when a pass adds or removes values.

Add the hidden `-instnamer-after-each-pass` option to name unnamed
arguments,
basic blocks, and non-void instructions before the first snapshot and
after
each new pass manager pass. It reuses InstructionNamer and keeps a
counter
across the pipeline, so a generated name is not reused after its value
is
removed. Existing names are preserved.

This makes text snapshots easier to compare, but names do not identify
instruction objects: a pass can transfer or reuse an existing name. The
ordinary `instnamer` pass and `-print-changed=diff` output are
unchanged.
Implement the standard POSIX.1-2024 function `tcsetwinsize` in
`<termios.h>`.

Fixes #228380
Part of #228378

Implementation was assisted by Antigravity by analysing other functions
in header and reviewed by Aman Maurya.
…230446)

canReuseInstruction gave up once it had visited 16 values, counting
constants and values already known to be poison-contributors of S. These
values are not walked any further, so only charge the instructions we
have to analyze operands of.

This improves re-use across a number of workloads end-to-end:
dtcxzyw/llvm-opt-benchmark-nightly#1624

I noticed this while invesigating missed re-use after changing operand
order in SCEV (#230252)

Compile-time impact in the noise

https://llvm-compile-time-tracker.com/compare.php?from=10d4b33fe57fecd864b7f9dbdfa558ae2f1329ad&to=0a29e64eaf5e64e683a81529323428d9b111751d&stat=instructions:u

PR: #230446
…h LIS (#230441)

To update LiveIntervals, SplitCriticalEdge checked every virtual
register in the function for liveness at the end of the split block.
That made PHIElimination's edge splitting O(splits * vregs).

PHIElimination now computes the set of virtual registers live out of
each block before the first split and passes it through
SplitCriticalEdgeAnalyses. Only those registers are visited, and the set
for the new block is added once the intervals are updated.

This essentially resurrects the per-block sparse register sets used to
update LiveVariables, which were removed in #228618 and #230147, applied
to the LiveIntervals update instead.

Instructions retired in phi-node-elimination, x86_64 -O3, on a generated
chain of N compare blocks branching to a shared PHI block:

  N     before          after           after/before
  1k      224,562,067      94,012,986   0.42
  2k      862,552,029     351,548,735   0.41
  4k    3,279,201,709   1,199,726,325   0.37
  8k   12,769,826,390   4,296,440,567   0.34
  16k  50,117,918,793  16,061,464,876   0.32

gcc-c-torture compile/20001226-1.c (liveintervals,phi-node-elimination
on the pre-PHIElimination MIR): 3,274,667,934 -> 1,193,137,243.

Whole llc on the 16k input: 112,903,512,059 -> 78,878,832,446.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The AMX intrinsics that name tile registers (`llvm.x86.tileloadd64`,
`llvm.x86.tdpbssd` and the rest of the non-`_internal` forms) have no
memory attributes, so to the optimizer every call may read and write any
memory. A loop-invariant value that lives in memory is reloaded after
every tile instruction, a store isn't forwarded past a tile load, and
two reads of the same tile row aren't merged.

This PR models the tile registers and the tile configuration as
`target_mem0`, the way AArch64 models SME's ZT0 and ZA
(`target_mem0`/`target_mem1`; `SME_Load_Intrinsic` is
`[IntrRead<[ArgMem, ZA]>, IntrWrite<[ZA]>]`), and declares what each
intrinsic actually does:

| Intrinsics | Reads | Writes |
|---|---|---|
| `tileloadd64`, `tileloaddt164`, `tileloaddrs64`, `tileloaddrst164` |
pointer argument, tile state | tile state |
| `tilestored64` | tile state | pointer argument |
| `tilezero`, `tdp*`, `tcmm*` | tile state | tile state |
| AMX-AVX512 row reads (`tilemovrow*`, `tcvtrow*`) | tile state | none |
| `ldtilecfg` | pointer argument | tile state |
| `sttilecfg` | tile state | pointer argument |
| `tilerelease` | none | tile state |

Every one of them accesses `target_mem0`, so tile operations stay
ordered with each other (and with any call that doesn't declare its
memory effects), and the memory operations stay ordered with every
access their pointer may alias. The `_internal` intrinsics, which the
AMX configuration passes manage, are unchanged.

## Why

A tile loop usually reads its geometry through a pointer:

```c
for (size_t i = 0; i < n; i++)
    _tile_loadd(0, base + i * cfg->step, cfg->stride);
```

Today `cfg->step` and `cfg->stride` are reloaded, and the address
recomputed, after every `TILELOADD`. With this PR LICM hoists them
(`Transforms/LICM/X86/amx-memory-effects.ll`). The same applies to
values read from globals or through any pointer in a loop of tile loads
or dot products.

For scale, I measured the Rust equivalent (where I found this, via
`std::arch`'s unstable AMX intrinsics) on a Xeon 6975P-C: a tile-load
loop that reloads its invariants every iteration (19 instructions per
tile) runs about 10% slower from L2 than the same loop with them in
registers (10 instructions; 6.0 vs 5.45 ns per 1 KiB tile). From DRAM
the difference doesn't show.

## Backend

The intrinsics' new properties change what TableGen infers for the
instructions selected by pattern: the custom-inserter pseudos for
`TILEZERO`, the dot products and the row reads lose
`UnmodeledSideEffects` (the row reads also lose `MayStore`), and
`TILERELEASE` loses `UnmodeledSideEffects` and `MayLoad`. The pseudos
are expanded during instruction emission into instructions whose TMM
register operands carry the dependencies, and `TILERELEASE` keeps
`MayStore` and its implicit definitions of TMM0-7. No existing AMX
CodeGen test changes. If you'd rather keep the backend flags exactly as
they were, I can set `hasSideEffects = 1` on those definitions.

## Tests

- The first commit adds the tests with today's output; the second shows
the change:
- `Transforms/LICM/X86/amx-memory-effects.ll`: loads hoisted out of
loops of tile loads and dot products, not out of a loop of tile stores.
- `Transforms/EarlyCSE/X86/amx-tile-state.ll`: two row reads merge, but
not across a tile load or a dot product; a store is forwarded past a
tile load, not past a tile store.
- The 43 existing tests that use these intrinsics (`CodeGen/X86/AMX`,
`CodeGen/X86`, `Transforms/InstCombine/X86`, `Verifier`),
`Transforms/EarlyCSE/target-memory.ll` and
`TableGen/target-mem-intrinsic-attrs.td` pass unchanged. I built only
`llvm` locally, so I haven't run clang's or MLIR's AMX tests.

Assisted-by: Claude Opus 5.5
This patch extends cir.add/cir.fadd to work on matrixes of int/float,
and lower correctly to add/fadd in LLVMIR.

It also extends cir.vec.splat to work with a matrix as well, so we could
get mixed-matrix-int/float operations to work. I considered making this
its own operation, however it is so nearly identical to cir.vec.splat
(and will become more so as we extend matrix) that it didn't seem
valuable to consider it separately, particularly as they lower to the
same things in LLVM.

One limitation: Splat is sometimes constant-folded during simplify.
However, we don't yet have a constant attribute type for a matrix, so
this is left for future work.
…230251)

A multilib runtimes build sets <PROJECT>_LIBDIR_SUBDIR. flang-rt
libraries and libclc bitcode were installed with the same name as the
base library. This doesn't really work for `libclc` because the clang
driver doesn't look through multilibs currently, but it should still go
somewhere else so it doesn't clobber.
Codegen does not support all kinds of casts for these types.

This patch adds missing costmodel and codedgen tests.
The tentative resolution for CWG1432 was not consistent for the
resolution of CWG1395.

This fixes the crash reported in #228870.

Additionally, this update the status of related core issues and papers
touching the same wording. Clang was never affected because we never
fully implemented CWG1395.

Fixes #228870
Fixes #27357

Assisted-by: Opus 5.5
TypeSystemClike needs to know about the sizes of various language
builtins on the current target. The LanguageOpts class provides this
information by querying Clang. The implementation of
GetFloatTypeSemantics and GetBitIntByteSize mirrors the logic in
TypeSystemClang.

See also #225371

assisted-by: claude
Reverts #230163

See the comment
#230163 (review)

We want to go with another approach by custom lowering mask_beforefirst
for fixed length vectors in AArch64 which will make the early expansion
dead code.
Negative stride reverses only the group order, not the byte
order within a widened lane. Emit a positive-stride load plus
the reorder shuffle for this case.

Fixes #230398

Reviewers: 

Pull Request: #230496
…223662)

CompileUnit has several lazy members that can be accessed by several
threads but that don't synchronize their lazy-loading mechanism. This
causes that some threads see the in-flight values when they access these
values.

This patch synchronizes all lazy-loading using the module's mutex.

This is a prerequisite for PR #220395 that adds a basic multithreaded
test and which would otherwise randomly fail.
Implement the standard POSIX.1-2008 / POSIX.1-2024 function `waitid` in
`<sys/wait.h>`, bringing functions declared in `<sys/wait.h>` to 100%
completion.

Fixes #227798

### Notes
- Defines the POSIX `idtype_t` enumeration in `llvm-libc-types` along
with its proxy header.
- In `linux/sys-wait-macros.h`, undefines the kernel `P_*` macros
originating from `<linux/wait.h>` to prevent macro expansion over the
`idtype_t` enum constants.
- Adds unit tests in `libc/test/src/sys/wait/` and live process reaping
integration tests in `libc/test/integration/src/unistd/fork_test.cpp`.
Use olMemRegister for dataLock/Unlock operations.
Move OpenMP specific notifyDataMapped/UnMapped to libomptarget
implemented on top of olMemRegister.

Remove all associated interfaces in GenericPluginTy.

Assisted by Claude.
…maxu/minu reductions (#230332)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…lt instructions) (#229362)

Add support for FEAT_CFLT (Conditional Fault instructions), which
are optional from Armv9.7 onwards:
```
 CFLTEQ, CFLTNE,
 CFLTGT, CFLTLT,
 CFLTGE, CFLTLE,
 CFLTHI, CFLTLO,
 CFLTHS, CFLTLS,
 CFLTZ,  CFLTNZ,
 TFLTZ,  TFLTNZ,
 FLT.NE, FLT.EQ,
 FLT.MI, FLT.PL,
 FLT.VS, FLT.VC,
 FLT.HI, FLT.LS,
 FLT.GE, FLT.LT,
 FLT.GT, FLT.LE,
 FLT.AL, FLT.NV,
 FLT.CS, FLT.CC,
 FLT.HS, FLT.LO
```

---

<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
…tomic 64-byte load/store) (#229363)

Add support for FEAT_LSC64B (single-copy atomic 64-byte load/store)
instructions:

  - LDA64B
  - STL64B
  - STL64BV
  - STL64BV0

---

<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
andjo403 and others added 30 commits October 10, 2026 20:11
Constant fold `bitinsert` and `bitextract`.

- Bits of a `ConstantByte` are extracted or inserted, then reinterpreted
as if by `bitcast`.
- A `bitinsert` that overwrites every bit of the base becomes a
`bitcast` of the value.
- An out-of-range, `poison` source or `poison`/`undef` offset returns
`poison`.

```llvm
%a = bitextract i8, b32 512, i32 8      ; i8 2
%b = bitinsert b32 0, i8 1, i32 8       ; b32 256
%c = bitinsert b16 0, half 1.0, i32 0   ; b16 15360, the bits of half 1.0
%d = bitextract i8, b32 0, i32 25       ; i8 poison
```
The loop-aware instruction-count check scaled the per-lane count of
splat gather entries by the trip count, rejecting profitable trees.
A splat is a single broadcast regardless of the loop scale. Fixes
4-6% RAJAPerf Apps_DIFFUSION3DPA regression on AArch64 introduced
by #211096.

Reviewers: 

Pull Request: #230821
Only ConstantInt operands are negated, so drop nsw/nuw on
a non-constant-int constants.

Follow-up to #230453.

Reviewers: 

Pull Request: #230835
Do not use gather nodes for cast context evaluation to prevent a crash.

Reviewers: 

Pull Request: #230836
## Summary

Normalize positive subnormal inputs in `log2f` by converting the integer
significand to float exactly and accounting for the `2^-149` scale.
This replaces the existing floating-point multiplication by `2^23`.

Replacement prepared on main at
`681b57195f495c346c1d15229334bb0ffcbba957`.
The remaining diff changes only `libc/src/__support/math/log2f.h`.

The `log10f` table relocation is already covered upstream by #229173,
merged as `636bedc3b3c00266a95f139c4b79cb8b940c1a36`. It moved `LOG10_R`
and `LOG10_2` to namespace-scope inline constexpr storage. This update
preserves that implementation and drops our redundant table change,
leaving only the `log2f` subnormal optimization.

## Validation

- Clang 18 generic and native x86-64 builds passed 33,954,500
differential
comparisons each: 67,909,000 records total, with no mismatches in result
  bits, errno, or floating-point exception flags.
- Each build checks every positive subnormal in all four rounding modes,
plus explicit boundary/special values and deterministic sampled inputs.
A subset checks preservation of pre-existing errno and exception state.
- Fresh LLVM libc log2f/log10f smoke suites: 8/8 tests passed.

These are comparisons with the matching upstream implementation, not an
independent MPFR accuracy proof or cross-platform certification.

## Historical Performance

The earlier PR reported 49.68 ns saved for the log2f denormal range.
That measurement predates this rebase; it has not been rerun on this
base.
No new performance claim is made for this update. The earlier log10f
timings no longer describe a change in this PR.

Signed-off-by: sriramshastry <sriramshastry@gmail.com>
Failed shuffle checks left a partial mask that adjustExtracts still
consumed. Lanes absent from the mask (out-of-bounds, poison, undef
index) also widened the shuffle, so valid two-vector gathers were
rejected.

Reviewers: 

Pull Request: #230847
Options read with getNumOccurrences() become OptionalValueField,
OptionalEnumField, or OptionalBoolField; readers apply the previous
fallback with value_or()/valueOr(). RecordStackHistoryMode moves to
InstrumentationOptions.h as an enum class. -profile-correlate no longer
accepts an empty value.

The cl::lists become ListFields. The string lists, which were not
cl::CommaSeparated, now split each value on commas too; no in-tree user
passes a comma.

-no-pgo-warn-mismatch (written by LTOBackend) and
-hwasan-mapping-offset{,-dynamic} (ordered by getPosition()) stay
cl::opt.

Aided by Opus 5.5
#230861)

`ppc32_elf_reloc.s`, introduced in #229933, fails on Windows.

```
Expression 'decode_operand(rel16_lo, 2) = (object - rel16_ha) [15:0]' is false: 0xfffffffffffffff4 != 0xfff4
```

`decode_operand` sign-extends the left side, so `0xfff4` becomes
`0xfffffffffffffff4`. The slice on the right-hand side stays `0xfff4`
and the test fails.

The fix is to slice the left side:

```diff
 # R_PPC_REL16_HA and R_PPC_REL16_LO
-# rtdyld-check: decode_operand(rel16_ha, 2) = (object - rel16_ha + 0x8000) [31:16]
+# rtdyld-check: decode_operand(rel16_ha, 2)[15:0] = (object - rel16_ha + 0x8000)[31:16]
 rel16_ha:
    addis 3, 3, object-rel16_ha@ha
-# rtdyld-check: decode_operand(rel16_lo, 2) = (object - rel16_ha) [15:0]
+# rtdyld-check: decode_operand(rel16_lo, 2)[15:0] = (object - rel16_ha)[15:0]
 rel16_lo:
    addi 3, 3, object-rel16_ha@l
```

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Truncate the source of a `uint_to_fp` when it is known to fit in a
narrower
type the target can convert from directly.
For example:
```
    uitofp (and i64 %x, 255) to float
```
On AMDGPU this becomes a single `v_cvt_f32_ubyte0` instead of the
generic
i64 to f32 expansion.
…upported under ASan (#230870)

These tests are flaky on x86_64 Linux ASan bots with `out of range of
Delta32
fixup` JITLink errors, similar to #102858, #135401, and #150242.

This likely happens because `InProcessMemoryManager` maps each
incremental
module with a separate `mmap` call, and depending on the address space
layout
some allocations appear to end up on opposite sides of ASan's large
allocator
reservation (> 2 GiB apart).

Assisted-by: Gemini
…229011)

lowerDYNAMIC_STACKALLOC wrapped the __ve_grow_stack call and the
GETSTACKTOP stack-pointer read in a zero-sized CALLSEQ_START/CALLSEQ_END
pair. The call it contains emits its own CALLSEQ, so the outer bracket
only produced a nested ADJCALLSTACKDOWN 0 / ADJCALLSTACKUP 0 around the
inner ADJCALLSTACKDOWN / ADJCALLSTACKUP which is illegal.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply@anthropic.com>
The LLVM libc FreeBSD CI has recently been failing because upstream 
package repositories require a newer FreeBSD version than the pinned 
15.0 image.

This patch updates the FreeBSD version specification from 15.0 to 15, 
allowing the action to automatically pull the latest minor snapshot. 
It also bumps the VM action to the latest release for improved 
compatibility with Ubuntu 26.04 runners.

ref:
https://github.com/llvm/llvm-project/actions/runs/37893910030/job/113701458175?pr=230371

- before

```
  Processing entries: 
  Newer FreeBSD version for package zh-qe:
  To ignore this error set IGNORE_OSVERSION=yes
  - package: 1501000
  - running userland: 1500068
 ```

- after

```
  [2026-10-10 10:29:20.317] Using release: 15.1
[2026-10-10 10:29:20.318] Downloading
https://github.com/anyvm-org/freebsd-builder/releases/download/v2.2.8/freebsd-15.1.qcow2.zst
```
Existing deserializers of SymbolLookupSet were spelling out the SPS type
in full (SPSSequence<SPSTuple<SPSString, bool>>). Define an
SPSSymbolLookupSet typedef and use in instead so that deserialization
points can pick up any future changes automatically.
Model the lane with the absorbing constant (0 for mul/and, -1 for or) of
a copyable node as op(V, undef), so the operand column of the other
lanes gets an undef lane. Cover such undef lanes in a column of
consecutive loads with a single frozen vector load, if the whole range
is dereferenceable.

Fixes #46897

Assisted-by: Cursor

Reviewers: RKSimon

Pull Request: #228872
Existing serializers of SymbolLookupResult were spelling out the SPS
type in full (SPSSequence<SPSOptional<SPSExecutorAddr>>). Define an
SPSSymbolLookupResult typedef and use in instead so that serialization
points can pick up any future changes automatically.

This is the result-side counterpart to 8ff4f38, which added a
typedef for SymbolLookupSet.
Fixes issue introduced by #181918 on gfx10+ where an immediate can get
commuted from src0 to src1 but then fail to get commuted back to src0
due to the legality checks in `isLegalToSwap`.

This PR relaxes the checks in `isLegalToSwap`, since gfx10+ allows the
immediate to be in locations other than src0. Relaxing these checks
causes MachineCSE to also successfully commute immediate operands out of
src0, which is the reason behind all the lit tests that required
modification. The PR also adds 2 new tests.

Co-authored by: Claude Code

Fixes: LCOMPILER-2920

---------

Co-authored-by: Claude <noreply@anthropic.com>
This PR restores VOPD pair formation after the legality checks
relaxation introduced by #229906. Before the legality check relaxation,
MachineCSE was commuting immediate operands from src1 to src0, and then
failing to commute them back, which inadvertently results in the
immediate operands in src0, and the VOPD pairing would succeed. After
the legality relaxation, MachineCSE is now able to successfully commute
the immediates back from src0 to src1, which breaks VOPD pairing since
the pass expected the immediates to be in src0 position. This change
adds a check in GCNVOPDUtils.cpp which checks if a commute is necessary
to allow the VOPD pairing, and then records that finding so that
GCNCreateVOPD applies the commute before creating the VOPD pair.

Co-authored by: Claude Code

---------

Co-authored-by: Claude <noreply@anthropic.com>
lowerLOADi1() rewrites an i1 load into a zext load to i16 plus a
truncate, and returns the (value, chain) pair as a MERGE_VALUES node.

LegalizeLoadOps installs that pair with

  RChain = Res.getValue(1);
  DAG.ReplaceAllUsesOfValueWith(SDValue(Node, 1), RChain);

so the second value of the MERGE_VALUES becomes the replacement for the
original load's chain result. Returning LD->getChain() therefore rewires
every memory operation that followed the original load to that load's
predecessor, and leaves the new zext load's chain result with no users.
The ordering edge between the new load and those memory operations is
dropped, so nothing in the DAG keeps them in order beyond whatever data
dependency happens to exist between them.

Return newLD.getValue(1) instead, so the edge is preserved.
In the common case where LockT = std::scoped_lock<std::mutex> is the
desired lock (and mutex) type, this allows us to write:

  LockedAccess<T> getValue() { return { Value, Mutex }; }

without having to spell out the type for LockT.
…anaged memory (#229213)

With -gpu=mem:managed, descriptors can live in managed memory, so a
descriptor built on the device may be read on the host. The type
descriptor address in its addendum then pointed at the device copy of
the type info, and host code dereferencing it crashed.

Make the host copy of the type info the single shared copy:

- Add the cuf-shared-type-info pass. It makes host type-info globals
  writable, places them in the __nv_type_info section, and drops
  acc.declare from type info on both host and device so that OpenACC
  declare constructors no longer copy it to the device. For each type
  descriptor used in the GPU module, it creates a managed pointer
  global <dt>Xhostaddr<tag>. The tag is a per-unit hash, which keeps
  the name unique when each unit has its own device module. The GPU
  module gets a cuf.shared_type_descs dictionary mapping each type
  descriptor to its pointer.
- CUFAddConstructor: add the cuda-managed-type-info option. It
  registers the __nv_type_info section pages with cudaHostRegister.
  After CUFInitModule, the constructor stores the host type descriptor
  address into each managed pointer.
- CodeGen: in GPU functions, load type descriptor addresses through the
  managed pointer for embox/rebox, fir.type_desc and fir.address_of
  outside global initializers.
- Runtime: add CUFRegisterHostMemoryRange. It page-aligns the range,
  registers it with cudaHostRegister(Portable|Mapped), and treats
  "already registered" as success.

Type-bound procedure calls and user-defined assignment or finalization
from device code are not supported. Type info from units compiled
without the option is not shared.
GFX13 widens the barrier member count in M0 to 8 bits.
…0883)

For coverage and PGO, `clang -mllvm
-profile-correlate={debug-info,binary}` moves the per-function profile
metadata out of the memory image at runtime.

After the cl::opt to TableGen migration (#230746),
InstrumentationOptions.h includes InstrProfCorrelator.h only for this
enum, adding about 0.25s to each Instrumentation TU. Fix the compile
time regression by moving the type to its own header as a scoped enum.

LLM-aided
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.