Skip to content

Inlining + target_feature broken in powerpc64 #60637

Description

@gnzlbg

godbolt

#![feature(repr_simd, powerpc_target_feature)]
#![allow(non_camel_case_types)]

#[repr(simd)] pub struct u32x4(u32, u32, u32, u32);

impl u32x4 {
    #[inline]
    // #[inline(always)]
    fn splat(x: u32) -> Self {
        u32x4(x, x, x, x)
    }
}

#[target_feature(enable = "altivec")]
pub unsafe fn splat_u32x4(x: u32) -> u32x4 {
    u32x4::splat(x)
}

with #[inline] that code produces a function call within splat_u32x4 (b example::u32x4::splat) to u32x4::splat, which is not eliminated, even though this method is module private. With #[inline(always)], u32x4::splat is inlined into splat_u32x4, and no code for u32x4::splat is generated.

#[inline] should not be needed here, much less #[inline(always)], yet without #[inline(always)] this produces bad codegen.

Removing the #[target_feature] attribute from splat_u32x4 fixes the issue, no #[inline] necessary: godbolt. So there must be some interaction between inlining and target features going on here.

Activity

  1. added
    C-bugCategory: This is a bug.
    O-PowerPCTarget: PowerPC processors
    T-compilerRelevant to the compiler team, which will review and decide on the PR/issue.
    on May 8, 2019
  2. added
    A-SIMDArea: SIMD (Single Instruction Multiple Data)
    on Nov 20, 2023
  3. sayantn commented on Jun 29, 2024

    @sayantn
    Contributor

    This is something on the LLVM side. See this comment. This happens in all microarchitectures except for ARM and x86 I think.

  4. added
    A-LLVMArea: Code generation parts specific to LLVM. Both correctness bugs and optimization-related issues.
    on Jun 29, 2024
  5. sayantn commented on Apr 13, 2025

    @sayantn
    Contributor
  6. 2 remaining items

  7. added a commit that references this issue on Jan 24, 2026
  8. added 4 commits that reference this issue on Jan 24, 2026
  9. added a commit that references this issue on Jan 24, 2026
  10. added a commit that references this issue on Jan 25, 2026
  11. folkertdev commented on Jan 27, 2026

    @folkertdev
    Contributor

    The final macos issue was resolved by

    and now produces

    mov rbp, rsp
    pavgw xmm0, xmm1
    
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    A-LLVMArea: Code generation parts specific to LLVM. Both correctness bugs and optimization-related issues.A-SIMDArea: SIMD (Single Instruction Multiple Data)C-bugCategory: This is a bug.O-PowerPCTarget: PowerPC processorsT-compilerRelevant to the compiler team, which will review and decide on the PR/issue.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions