How do you use <code>__m256d</code>? Say I want to use the Intel AVX instruction <code>_mm256_add_pd</code> on a simple <code>Vector3</code> class with 3-64 bit <code>double</code> precision components (<code>x</code>, <code>y</code>, and <code>z</code>). What is the correct way to use this? Since <code>x</code>, <code>y</code> and <code>z</code> are members of the <code>Vector3</code> class, _can I declare them in <code>union</code> with an <code>__m256d</code> variable? <pre class="prettyprint"><code>union Vector3 { struct { double x,y,z ; } ; __m256d _register ; // the Intel register? } ; </code></pre> Then can I go: <pre class="prettyprint"><code>Vector3 add( const Vector3& o ) { Vector3 result; result._register = _mm256_add_pd( _register, o._register ) ; // add 'em return result; } </code></pre> Is that going to work? Or do I need to declare temporaries, <pre class="prettyprint"><code>Vector3 add( const Vector3& o ) { __m256d d1 = *(__m256d*)(&x) ; // ? Cast to __m256d? __m256d d2 = *(__m256d*)(&o.x) ; // ? Cast to __m256d? __m256d result = _mm256_add_pd( d1, d2 ) ; // add 'em return Vector3( result ) ; // make a ctor that accepts __m256d? } </code></pre> <h3>Edit</h3> I came up with this example, <pre class="prettyprint"><code>#include <stdio.h> #include <intrin.h> int main() { __m256d a, b, res; for( int i = 0; i < sizeof(__m256d)/sizeof(double); i++ ) { a.m256d_f64[i] = i ; b.m256d_f64[i] = 2*i ; } // Perform __4__ adds. res = _mm256_add_pd(a, b); for( int i = 0; i < sizeof(__m256d)/sizeof(double); i++ ) { printf("%f + %f = %f\n", a.m256d_f64[i], b.m256d_f64[i], res.m256d_f64[i]); } puts(""); } </code></pre> I guess the question is now, does <code>_mm256_add_pd</code> do load operations automatically, or will something get messed up if I don't declare my <code>__m256d</code> registers as locals close to where they are used? (I'm afraid of a hotel room / deskdrawer type problem) <h3>Edit 2:</h3> I tried adding an <code>__m256</code> register to my rather large project, and I got a whole bunch of <blockquote> error C2719: 'value': formal parameter with __declspec(align('32')) won't be aligned </blockquote> Errors, it leads me to believe that you can't keep <code>__m256</code> registers inside a class, instead they should be declared as locals?

First I'd like to clear up a little confusion. <code>__m256d</code> isn't a type of register, it's a data type that can be loaded into an AVX register. A <code>__m256d</code> is no more a register than an <code>int</code> is a register. There are a few ways to get data in and out of an <code>__m256d</code> (or any other vector type): Using a <code>union</code>: Yes, the <code>union</code> trick works. It works very well, since the union will generally have the correct alignment (although <code>malloc</code> might not, use <code>posix_memalign</code> or <code>_aligned_malloc</code>). <pre class="prettyprint"><code>class Vector3 { public: Vector3(double xx, double yy, double zz); Vector3(__m256d vvec); Vector3 operator+(const Vector3 &other) const { return Vector3(_mm256_add_pd(vec, other.vec)); } union { struct { double x, y, z; }; __m256d vec; // a data field, maybe a register, maybe not }; }; </code></pre> Using intrinsics: Inside a function, it's usually easier to use intrinsics to get data in and out of a vector type. <pre class="prettyprint"><code>__m256d vec = ...; double x, y, z; vec = _mm256_add_pd(vec, _mm256_set_pd(x, y, z, 0.0)); </code></pre> Using pointer casts: Casting pointers is the last resort for a couple of reasons. <ol> <li>The pointer might not be aligned correctly.</li> <li>Casting pointers can sometimes mess with the compiler's aliasing analysis.</li> <li>Pointer casting bypasses a number of safety guarantees.</li> </ol> So I'd only use pointer casting to plow through a big array of data.

Using __m256d registers

Tags:

c++

x86

avx

intel

simd

How do you use __m256d?

Say I want to use the Intel AVX instruction _mm256_add_pd on a simple Vector3 class with 3-64 bit double precision components (x, y, and z). What is the correct way to use this?

Since x, y and z are members of the Vector3 class, _can I declare them in union with an __m256d variable?

union Vector3
{
  struct { double x,y,z ; } ;
  __m256d _register ;  // the Intel register?
} ;

Then can I go:

Vector3 add( const Vector3& o )
{
  Vector3 result;
  result._register = _mm256_add_pd( _register, o._register ) ; // add 'em
  return result; 
}

Is that going to work? Or do I need to declare temporaries,

Vector3 add( const Vector3& o )
{
  __m256d d1 = *(__m256d*)(&x) ; // ? Cast to __m256d?
  __m256d d2 = *(__m256d*)(&o.x) ; // ? Cast to __m256d?
  __m256d result = _mm256_add_pd( d1, d2 ) ; // add 'em
  return Vector3( result ) ; // make a ctor that accepts __m256d?
}

Edit

I came up with this example,

#include <stdio.h>
#include <intrin.h>

int main()
{
  __m256d a, b, res;

  for( int i = 0; i < sizeof(__m256d)/sizeof(double); i++ )
  {
    a.m256d_f64[i] = i ;
    b.m256d_f64[i] = 2*i ;
  }

  // Perform __4__ adds.
  res = _mm256_add_pd(a, b);

  for( int i = 0; i < sizeof(__m256d)/sizeof(double); i++ )
  {
    printf("%f + %f = %f\n", a.m256d_f64[i], b.m256d_f64[i], res.m256d_f64[i]);
  }
  puts("");
}

I guess the question is now, does _mm256_add_pd do load operations automatically, or will something get messed up if I don't declare my __m256d registers as locals close to where they are used? (I'm afraid of a hotel room / deskdrawer type problem)

Edit 2:

I tried adding an __m256 register to my rather large project, and I got a whole bunch of

error C2719: 'value': formal parameter with __declspec(align('32')) won't be aligned

Errors, it leads me to believe that you can't keep __m256 registers inside a class, instead they should be declared as locals?

346

asked Oct 13 '12 17:10

bobobobo

1 Answers

First I'd like to clear up a little confusion. __m256d isn't a type of register, it's a data type that can be loaded into an AVX register. A __m256d is no more a register than an int is a register. There are a few ways to get data in and out of an __m256d (or any other vector type):

Using a union: Yes, the union trick works. It works very well, since the union will generally have the correct alignment (although malloc might not, use posix_memalign or _aligned_malloc).

class Vector3 {
public:
    Vector3(double xx, double yy, double zz);
    Vector3(__m256d vvec);


    Vector3 operator+(const Vector3 &other) const
    {
        return Vector3(_mm256_add_pd(vec, other.vec));
    }

    union {
        struct {
            double x, y, z;
        };
        __m256d vec; // a data field, maybe a register, maybe not
    };
};

Using intrinsics: Inside a function, it's usually easier to use intrinsics to get data in and out of a vector type.

__m256d vec = ...;
double x, y, z;
vec = _mm256_add_pd(vec, _mm256_set_pd(x, y, z, 0.0));

Using pointer casts: Casting pointers is the last resort for a couple of reasons.

The pointer might not be aligned correctly.
Casting pointers can sometimes mess with the compiler's aliasing analysis.
Pointer casting bypasses a number of safety guarantees.

So I'd only use pointer casting to plow through a big array of data.

152

answered Sep 18 '22 11:09

Dietrich Epp

Related questions
                            
                                Get the list of methods of a class
                            
                                enable_shared_from_this and objects on stack
                            
                                Finding "~/Library/Application Support" from C++?
                            
                                What book would cover theory for 3D game development mathematics? [closed]
                            
                                "noexcept" vs "Throws: nothing" [closed]
                            
                                Configure gtest to show failed test only in console
                            
                                Once an array-of-T has decayed into a pointer-to-T, can it ever be made into an array-of-T again?
                            
                                Is it safe to cast arbitrary values of the underlying type to a strongly-typed enum type?
                            
                                How to approximate the count of distinct values in an array in a single pass through it
                            
                                Detecting the parameter types in a Spirit semantic action
                            
                                qtextedit - resize to fit
                            
                                constexpr, static_assert, and inlining
                            
                                Passing a C# class object in and out of a C++ DLL class
                            
                                How do I profile a MEX-function in Matlab
                            
                                is c++ Template Metaprogramming a form of functional programming
                            
                                Two-way C++ communication over serial connection
                            
                                Why is Visual C++ not performing return-value optimization on the most trivial code?
                            
                                Clang with -faddress-sanitizer on Windows
                            
                                Converting float vector to 16-bit int without saturating
                            
                                Random sequence iteration in O(1) memory?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With