4 - Approximations in Scientific Computing; Computer Arithmetic

Updated 4 Oct 2026

Overview

  • Dealing with real numbers involves approximations/inexactness
  • Approximation/inexactness occurs both before and during computations

เวลาเราทำงานกับ จำนวนจริง บนคอมพิวเตอร์หรือในวิชาคณิต มันไม่มีทางเป๊ะ 100% หรอก จะมีความ “คร่าว ๆ” (approximation) แอบอยู่เสมอ ทั้งจาก ก่อนเริ่มคำนวณ และ ระหว่างคำนวณ เชื่อมั้ยล่ะ เลยลร่ะ

Sources of Approximations Before Computation Begins

1. Modeling

  • Some physical features of the problem or system under study may be simplified or omitted
  • Examples: friction, viscosity, air resistance

โลกจริงมันซับซ้อนเกินไป เราเลยต้องตัดทอนบางอย่างออก เช่น จำลองรถแต่ไม่ใส่แรงเสียดทาน → คำตอบมันก็จะไม่ตรงเป๊ะกับโลกจริง

2. Empirical Measurement

  • Instruments for measurement have finite precision
  • Small sample size
  • Random noise during measurement

เครื่องมือวัดมีความละเอียดจำกัด + อาจมีสัญญาณรบกวน + ข้อมูลน้อยเกินไป → ค่าที่ได้จึงไม่ใช่ค่าจริงเป๊ะ ๆ

3. Previous Computations

  • Input may be produced by a previous computation, whose results were only approximate

ถ้า input มาจากการคำนวณรอบก่อน ซึ่งตอนนั้นก็ไม่เป๊ะ 100% อยู่แล้ว → พอเอามาใช้ต่อ ความคลาดเคลื่อนก็ถูกส่งต่อมา เออเริ่ด

Sources of Approximations During Computation

1. Truncation or Discretization

  • Using a finite number of terms in an infinite series
  • Example: ex=1+x1!+x22!+x33!+⋯e^x = 1 + \frac{x}{1!} + \frac{x^2}{2!} + \frac{x^3}{3!} + \cdots
    • อันนี้ก็ใช้ Taylor Series มา Approximate!
    • สูตรหลายอันใช้อนุกรมอนันต์ เช่น exe^x แต่เราเอามาใช้แค่ไม่กี่พจน์ → เลยไม่ตรงจริง
  • Replacing derivatives by finite differences (which is only approximation of true derivatives)

2. Rounding

  • Floating-point variables can store finite amount of precision
  • Results of arithmetic operations must be rounded
  • Exact arithmetic is possible but very slow (cannot be used for large problems)

ตัวเลขจริงมันมีทศนิยมยาวไม่สิ้นสุด แต่คอมพิวเตอร์เก็บได้แค่บางหลัก → ต้องปัด

Error Amplification

  • While each of these approximations may be small, the error may be amplified by the nature of the problem being solved or the algorithm being used, or both

แม้แต่ error เล็ก ๆ ก็สามารถถูกขยายโดยตัวโจทย์หรือ algorithm ได้ เช่น บางสมการไวต่อการเปลี่ยนค่ามาก → ความผิดพลาดยิ่งทวีคูณ ถูกเลยล่ะ

Types of Error

Absolute Error and Relative Error

  • Absolute error = ∣|approximate value −- true value∣|
    • ส่วนต่างตรง ๆ ระหว่างค่าที่ได้กับค่าจริง
  • Relative error = absolute errortrue value\boxed{\frac{\text{absolute error}}{\text{true value}}}
    • ส่วนต่างนั้น เทียบกับขนาดของค่าจริง (ดูว่า “ผิดไปกี่ %”)

For scalar values:

  • If x^∈R\hat{x} \in \mathbb{R} is an approximation to x∈Rx \in \mathbb{R}, then the relative error is: ∣x^−x∣∣x∣\boxed{\frac{|\hat{x} - x|}{|x|}}
    For vectors:
  • If x^∈Rn\hat{x} \in \mathbb{R}^n is an approximation to x∈Rnx \in \mathbb{R}^n, then the relative error is: ∥x^−x∥∥x∥\boxed{\frac{\|\hat{x} - x\|}{\|x\|}}

Two Types of Computational Errors

Computational error (or error made during a computation) has two types:

1. Truncation Error

  • The difference between the true result and the result that would be produced by a given algorithm using exact arithmetic
  • The error from truncating an infinite series, replacing derivatives by finite differences, etc.

2. Rounding Error

  • The difference between the result produced by a given algorithm using exact arithmetic and the result produced by the same algorithm using floating-point arithmetic

Total computational error = truncation error + rounding error\large{\text{Total computational error = truncation error + rounding error}}

Floating Point Arithmetic

Floating-Point Number System

  • คอมพิวเตอร์เก็บเลขแบบ Scientific notation base 2
  • Floating-point number system: ±(d0.d1d2…dp−1)2×2E\boxed{\pm(d_0.d_1d_2 \ldots d_{p-1})_2 \times 2^E}
  • pp = number of mantissa digits (precision)
  • EE = integer exponent (L≤E≤UL \leq E \leq U)
  • did_i = either 0 or 1

Examples:

  • 3=(1.1)2×213 = (1.1)_2 \times 2^1
  • 0.5=(1.0)2×2−10.5 = (1.0)_2 \times 2^{-1}
  • 10=(1.01)2×2310 = (1.01)_2 \times 2^3

IEEE Standard Double-Precision Floating-Point Systems

  • Example: the double variable in JAVA

  • p=53p = 53, L=−1022L = -1022, U=1023U = 1023

  • Many minor adjustments:

  • d0d_0 is assumed to always be 1 (so it is not stored)

  • Using bias of 1023 for the exponent

  • Special values for 0, infinity, and NaN (Not-a-Number)

  • เก็บ 64 บิต: 1 บิตบอกเครื่องหมาย, 11 บิตเก็บเลขชี้กำลัง, 52 บิตเก็บทศนิยม

  • ใช้ bias 1023 สำหรับ exponent (คือเก็บค่า exponent บวก 1023 ไว้จริง ๆ)

  • มีค่าพิเศษ เช่น 0, infinity, NaN

Bias of 1023 for the exponent:

  • (00000000001)2=−1022(00000000001)_2 = -1022
  • (00000000010)2=−1021(00000000010)_2 = -1021
  • ⋮\vdots
  • (11111111110)2=1023(11111111110)_2 = 1023
  • (00000000000)2(00000000000)_2 and (11111111111)2(11111111111)_2 are reserved for special values

Memory Layout:

  • One double variable is 64 bits:
  • 1 bit for sign
  • 11 bits for the exponent
  • 52 bits for the mantissa

Unit Roundoff (Machine Epsilon)

  • ความหมาย: คือ “ความผิดพลาดสัมพัทธ์ที่เล็กที่สุด” ที่คอมพิวเตอร์หลีกเลี่ยงไม่ได้ในการเก็บหรือคำนวณตัวเลข 1 ครั้ง
  • เหมือนบอกว่า ต่อให้เทพแค่ไหน ก็เลี่ยงการปัดเศษไม่ได้เลย

Unit roundoff or machine epsilon is the maximum possible relative error resulting from one scalar operation in floating-point arithmetic.

Denoted as ϵmach\epsilon_{\text{mach}}.

If digits are simply chopped off:

ϵmach=21−p\boxed{\epsilon_{\text{mach}} = 2^{1-p}}

If rounding to nearest (in case of tie, round to number whose last stored digit is 0):

ϵmach=2−p\boxed{\epsilon_{\text{mach}} = 2^{-p}}
For IEEE double-type variables: ϵmach=2−52≈2.2×10−16\boxed{\epsilon_{\text{mach}} = 2^{-52} \approx 2.2 \times 10^{-16}}

แปลว่า เวลาใช้ double ใน Java, C++, Swift อะไรพวกนี้ แต่ละ operation จะมี error ได้ระดับ 10−1610^{-16} เลยทีเดียว (เล็กมาก แต่ไม่เป็นศูนย์)

Examples of Roundoff Error

Example 1 (บวกเลขแล้วเก็บไม่ได้)

Assume chopping-off-extra-digits "rounding" with p=2p = 2:

  • Let x=(1.0)2×22=4x = (1.0)_2 \times 2^2 = 4 and y=(1.0)2×20=1y = (1.0)_2 \times 2^0 = 1
  • z=x+y=5=(1.01)2×22z = x + y = 5 = (1.01)_2 \times 2^2
  • But zz cannot be stored exactly in floating-point arithmetic (too many digits)
  • So we have to round zz to (1.0)2×22=4(1.0)_2 \times 2^2 = 4

Example 2 (ลบแล้วหล่น)

Suppose p=3p = 3:

  • Let x=(1.10)2×23=12x = (1.10)_2 \times 2^3 = 12 and y=(1.10)2×21=3y = (1.10)_2 \times 2^1 = 3
  • z=x−y=9=(1.001)2×23z = x - y = 9 = (1.001)_2 \times 2^3
  • Again, zz cannot be stored exactly
  • So we have to round zz to (1.00)2×23=8(1.00)_2 \times 2^3 = 8

Example 3 (คูณแล้วหาย)

Suppose p=2p = 2:

  • Let x=y=3=(1.1)2×21x = y = 3 = (1.1)_2 \times 2^1
  • z=xy=9=(1.001)2×23z = xy = 9 = (1.001)_2 \times 2^3
  • zz cannot be stored exactly
  • So we have to round zz to (1.0)2×23=8(1.0)_2 \times 2^3 = 8

Catastrophic Cancellation

อันนี้คือ ตัวร้ายในวงการ numerical

Definition

Consider subtracting two numbers having the same sign and similar magnitude such that the result has much smaller magnitude.

Example

  • Say, x^=150422\hat{x} = 150422 and y^=150419\hat{y} = 150419
  • We want to compute: z^:=x^−y^=150422−150419=3\hat{z} := \hat{x} - \hat{y} = 150422 - 150419 = 3
  • But x^\hat{x} and y^\hat{y} are likely to be approximations (from previous computation or measurement)
  • Suppose x^\hat{x} is the approximation of x=150421x = 150421
  • And y^\hat{y} is the approximation of y=150420y = 150420
  • The true difference is: z=x−y=150421−150420=1z = x - y = 150421 - 150420 = 1

Error Analysis:

Relative error of the output:
∣z^−z∣∣z∣=∣3−1∣∣1∣=2\frac{|\hat{z} - z|}{|z|} = \frac{|3 - 1|}{|1|} = 2
Relative errors of the inputs:
∣x^−x∣∣x∣=∣150422−150421∣∣150421∣=1150421\frac{|\hat{x} - x|}{|x|} = \frac{|150422 - 150421|}{|150421|} = \frac{1}{150421}
∣y^−y∣∣y∣=∣150419−150420∣∣150420∣=1150420\frac{|\hat{y} - y|}{|y|} = \frac{|150419 - 150420|}{|150420|} = \frac{1}{150420}

The input relative errors are much smaller than the output relative error!

General Explanation

  • Subtracting two numbers having the same sign and similar magnitude such that the result has much smaller magnitude can result in very large error
  • This is because the absolute error of the result is about the same magnitude as the absolute errors of the two numbers
  • But the true value of the result is much smaller in magnitude than the true values of the two numbers
  • Hence we have very high relative error in the result
  • Multiplication, division, and addition of two numbers having the same sign are "safe"

Important Note

If you have catastrophic cancellation in the middle of your computation, your final result may or may not have large relative error (it depends on subsequent operations).

Examples:

  • 1/x1/x has about the same relative error as that of xx

  • If xx is about 1, then x+1000x + 1000 has much smaller relative error than that of xx

  • Catastrophic cancellation ในขั้นตอนกลาง ไม่ได้แปลว่า final result จะพังเสมอไป → ขึ้นกับว่าเราเอาค่าที่ผิดนั้นไปทำอะไรต่อ

  • เช่น 1/x1/x → error ของมันจะใกล้เคียงกับ error ของ xx

  • หรือ x+1000x+1000 (ถ้า xx ประมาณ 1) → error เล็กมากจนแทบไม่ส่งผล

Summary

This material covers the fundamental sources of error in scientific computing:

  1. Pre-computation errors: modeling simplifications, measurement limitations, and accumulated errors from previous calculations
  2. Computational errors: truncation errors from algorithmic approximations and rounding errors from finite precision arithmetic
  3. Floating-point representation: IEEE double precision standard with 64-bit representation and machine epsilon of approximately 2.2×10−162.2 \times 10^{-16}
  4. Catastrophic cancellation: a critical numerical issue where subtracting nearly equal numbers can dramatically amplify relative error

Understanding these error sources is essential for developing robust numerical algorithms and interpreting computational results accurately.