04 Apriori Algorithm

Updated 4 Oct 2026

  • ถ้า itemset ใดเกิดไม่บ่อย superset ทั้งหมดของ itemset นั้น ก็จะเกิดไม่บ่อยด้วยเช่นกัน
  1. ตัด infrequent ออกไปก่อน

  2. Frequent items of size 1

    • {pasta}, {lemon}, {orange}, {cake}
  3. Frequent items of size 2

    • {pasta, lemon}, {pasta, orange}, {pasta, cake}
    • {lemon, orange}, {lemon, cake}
    • {orange, cake}
  4. something…

  5. Scan the database to calculate the support of remaining candidate itemsets of size 2

    • {pasta, lemon} = 3
    • {pasta, orange} = 3
    • {pasta, cake} = 2
    • {lemon, orange} = 2
    • {lemon, cake} = 1 (ELIMINATED)
    • {orange, cake} = 2
  6. Size 2

    • {pasta, lemon} = 3
    • {pasta, orange} = 3
    • {pasta, cake} = 2
    • {lemon, orange} = 2
    • {orange, cake} = 2
  7. Generate candidates of size 3 by combining frequent paris of itemsets

    • {pasta, lemon, orange}
    • {pasta, lemon, cake} → Infrequent (ELIMINATED)
    • {pasta, orange, cake}
    • {lemon, orange, cake} → Infrequent (ELIMINATED)

    Because {lemon, cake} is infrequent ที่เรา Eliminate ไปตอนแรกอะ เราก็สามารถ Cut ออกไปได้ (Using Property 2)

  8. Eliminate candidates of size 3 having a subset of size 2 that is infrequent

    • {pasta, lemon, orange}
    • {pasta, orange, cake}
  9. Count support

    • {pasta, lemon, orange} = 2
    • {pasta, orange, cake} = 2
  10. Generate candidates of size 4

    • {pasta, lemon, orange, cake}

APRIORI VS NAÏVE

  • The Apriori property can considerably reduce the number of itemsets to be considered.
  • In the previous example:
    • Naïve approach: 25-1 = 31 itemsets are considered.
    • By using the Apriori algorithm: 16 itemsets are considered.

ASSOCIATION RULES GENERATION

An association rule of the form A => B states that there is a correlation or association

between occurrences of the itemset A, known as the left hand side LHS or antecedent,

and the itemset B, known as the right hand side RHS or consequent.

  • อันนี้จะคำนึงด้วยอันไหนวางในตระกร้าก่อน มันก็มีผลนะ
Confidence(LHS=>RHS)=support(LHS∩RHS)support(LHS)Confidence(LHS= > RHS) =\frac{support(LHS\cap RHS)}{support(LHS)}