I'm reading this document: http://software.intel.com/en-us/articles/interactive-ray-tracing and I stumbled upon these three lines of code: <blockquote> The SIMD version is already quite a bit faster, but we can do better. Intel has added a fast 1/sqrt(x) function to the SSE2 instruction set. The only drawback is that its precision is limited. We need the precision, so we refine it using Newton-Rhapson: </blockquote> <pre class="prettyprint"><code> __m128 nr = _mm_rsqrt_ps( x ); __m128 muls = _mm_mul_ps( _mm_mul_ps( x, nr ), nr ); result = _mm_mul_ps( _mm_mul_ps( half, nr ), _mm_sub_ps( three, muls ) ); </code></pre> <blockquote> This code assumes the existence of a __m128 variable named 'half' (four times 0.5f) and a variable 'three' (four times 3.0f). </blockquote> I know how to use Newton Raphson to calculate a function's zero and I know how to use it to calculate the square root of a number but I just can't see how this code performs it. Can someone explain it to me please?

Given the Newton iteration <img src="https://i.stack.imgur.com/9JEAS.png" alt="y_n+1=y_n(3-x(y_n)^2)/2">, it should be quite straight forward to see this in the source code. <pre class="prettyprint"><code> __m128 nr = _mm_rsqrt_ps( x ); // The initial approximation y_0 __m128 muls = _mm_mul_ps( _mm_mul_ps( x, nr ), nr ); // muls = x*nr*nr == x(y_n)^2 result = _mm_mul_ps( _mm_sub_ps( three, muls ) // this is 3.0 - mul; /*multiplied by */ __mm_mul_ps(half,nr) // y_0 / 2 or y_0 * 0.5 ); </code></pre> And to be precise, this algorithm is for the inverse square root. Note that this still doesn't give fully a fully accurate result. <code>rsqrtps</code> with a NR iteration gives almost 23 bits of accuracy, vs. <code>sqrtps</code>'s 24 bits with correct rounding for the last bit. The limited accuracy is an issue if you want to truncate the result to integer. <code>(int)4.99999</code> is <code>4</code>. Also, watch out for the <code>x == 0.0</code> case if using <code>sqrt(x) ~= x * sqrt(x)</code>, because <code>0 * +Inf = NaN</code>.

To compute the inverse square root of <code>a</code>, Newton's method is applied to the equation <code>0=f(x)=a-x^(-2)</code> with derivative <code>f'(x)=2*x^(-3)</code> and thus the iteration step <pre class="prettyprint"><code>N(x) = x - f(x)/f'(x) = x - (a*x^3-x)/2 = x/2 * (3 - a*x^2) </code></pre> This division-free method has -- in contrast to the globally converging Heron's method -- a limited region of convergence, so you need an already good approximation of the inverse square root to get a better approximation.

Newton Raphson with SSE2 - can someone explain me these 3 lines

Tags:

c++

c

math

sse

newtons-method

I'm reading this document: http://software.intel.com/en-us/articles/interactive-ray-tracing

and I stumbled upon these three lines of code:

The SIMD version is already quite a bit faster, but we can do better. Intel has added a fast 1/sqrt(x) function to the SSE2 instruction set. The only drawback is that its precision is limited. We need the precision, so we refine it using Newton-Rhapson:

 __m128 nr = _mm_rsqrt_ps( x );   __m128 muls = _mm_mul_ps( _mm_mul_ps( x, nr ), nr );   result = _mm_mul_ps( _mm_mul_ps( half, nr ), _mm_sub_ps( three, muls ) );

This code assumes the existence of a __m128 variable named 'half' (four times 0.5f) and a variable 'three' (four times 3.0f).

I know how to use Newton Raphson to calculate a function's zero and I know how to use it to calculate the square root of a number but I just can't see how this code performs it.

Can someone explain it to me please?

790

asked Feb 07 '13 13:02

Marco A.

2 Answers

Given the Newton iteration y_n+1=y_n(3-x(y_n)^2)/2 , it should be quite straight forward to see this in the source code.

 __m128 nr   = _mm_rsqrt_ps( x );                  // The initial approximation y_0  __m128 muls = _mm_mul_ps( _mm_mul_ps( x, nr ), nr ); // muls = x*nr*nr == x(y_n)^2  result = _mm_mul_ps(                _mm_sub_ps( three, muls )    // this is 3.0 - mul;    /*multiplied by */ __mm_mul_ps(half,nr)  // y_0 / 2 or y_0 * 0.5  );

And to be precise, this algorithm is for the inverse square root.

Note that this still doesn't give fully a fully accurate result. rsqrtps with a NR iteration gives almost 23 bits of accuracy, vs. sqrtps's 24 bits with correct rounding for the last bit.

The limited accuracy is an issue if you want to truncate the result to integer. (int)4.99999 is 4. Also, watch out for the x == 0.0 case if using sqrt(x) ~= x * sqrt(x), because 0 * +Inf = NaN.

answered Oct 02 '22 06:10

Aki Suihkonen

To compute the inverse square root of a, Newton's method is applied to the equation 0=f(x)=a-x^(-2) with derivative f'(x)=2*x^(-3) and thus the iteration step

N(x) = x - f(x)/f'(x) = x - (a*x^3-x)/2       = x/2 * (3 - a*x^2)

This division-free method has -- in contrast to the globally converging Heron's method -- a limited region of convergence, so you need an already good approximation of the inverse square root to get a better approximation.

answered Oct 02 '22 07:10

Lutz Lehmann

Related questions
                            
                                Are two function pointers to the same function always equal?
                            
                                Structs vs classes in C++ [duplicate]
                            
                                Why does C++ linking use virtually no CPU?
                            
                                C++ nested classes accessibility
                            
                                Default initialization of C++ Member arrays?
                            
                                best way to do variant visitation with lambdas
                            
                                Qt foreach loop ordering vs. for loop for QList
                            
                                why is std::lock_guard not movable?
                            
                                Qt - add a hyperlink to a dialog
                            
                                Why define operator + or += outside a class, and how to do it properly?
                            
                                Simple object detection using OpenCV and machine learning
                            
                                Creating new types in C++
                            
                                How do I invoke the MinGW cross-compiler on Linux?
                            
                                Using std::tie as a range for loop target
                            
                                What are _mm_prefetch() locality hints?
                            
                                How can you detect if two regular expressions overlap in the strings they can match?
                            
                                How can i use tesseract ocr(or any other free ocr) in small c++ project?
                            
                                Should I use the same name for a member variable and a function parameter in C++?
                            
                                Boost::asio - how to interrupt a blocked tcp server thread?
                            
                                Are there any disadvantages to "multi-processor compilation" in Visual Studio?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With