Learn With Examples https://learnwithexamples.org/ Lets Learn things the Easy Way Tue, 04 Aug 2026 18:33:58 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.4 https://i0.wp.com/learnwithexamples.org/wp-content/uploads/2026/07/cropped-learnwithexamples-icon.png?fit=32%2C32&ssl=1 Learn With Examples https://learnwithexamples.org/ 32 32 228207193 How Electricity Works · Circuits Explained (Series vs Parallel) https://learnwithexamples.org/how-electricity-works/ https://learnwithexamples.org/how-electricity-works/#respond Tue, 04 Aug 2026 18:33:56 +0000 https://learnwithexamples.org/?p=865 Learn With Examples Why did one dead bulb kill a whole string of old fairy lights, while one dead lamp in your house changes nothing? Same electricity, two different wiring…

The post How Electricity Works · Circuits Explained (Series vs Parallel) appeared first on Learn With Examples.

]]>

Learn With Examples

Why did one dead bulb kill a whole string of old fairy lights, while one dead lamp in your house changes nothing? Same electricity, two different wiring choices — and that single difference explains most of the electrical world around you.

Reading time15 min
LevelBeginner friendly
Includes6 diagrams

Somewhere in a box in your attic there may still be a string of old Christmas lights with one job: to work. Half the time they didn’t. One bulb would fail and the entire string went dark, and you’d be left testing forty bulbs one at a time to find the culprit.

Meanwhile, in the same house, a bulb in the hallway could blow and nothing else noticed. The kitchen stayed lit. The fridge kept running.

That difference isn’t about the quality of the bulbs. It’s about how the wires connect them — in series, one after another in a single chain, or in parallel, each with its own path. Understanding that one distinction lets you explain fairy lights, house wiring, car headlights, torches, phone chargers, and why your kettle trips the breaker when the microwave is already running.

We’ll get there properly: first what electricity actually is, then what a circuit is, then the two ways of wiring one.

The whole article in four lines

Series and parallel, side by side

Series — one single path. The same current flows through everything, the voltage gets divided up between components, and if any one component fails, the path breaks and everything stops.

Parallel — multiple paths. Every branch gets the full voltage, the current splits between branches, and if one branch fails, the others carry on.

Your home is wired in parallel. That’s why every socket delivers the same voltage, why appliances work independently — and why adding more of them adds up current until a breaker steps in.

What electricity actually is

Inside a copper wire, the outer electrons of the copper atoms are only loosely attached and drift about randomly. Connect that wire to a battery and something changes: the random drifting gains a direction. All those electrons begin edging the same way, and that organised movement of charge is what we call an electric current.

Now for the fact that surprises almost everyone. Those electrons move remarkably slowly — a fraction of a millimetre per second in a typical household wire. Slower than a snail. An individual electron leaving your fuse box might take hours to reach the ceiling light.

So why does the light come on instantly? Because you’re not waiting for a particular electron to arrive. The wire is already full of them. Flipping the switch sends an electromagnetic push through the wire at close to the speed of light, and every electron along the way starts moving at once — the way water comes out of a tap immediately even though that particular water has been sitting in the pipe for hours.

The water analogy, and where it breaks

Electricity is invisible, which is why almost every explanation reaches for plumbing. It’s a genuinely good analogy, as long as you know its limits.

WATER CIRCUIT PUMP NARROWPIPE high pressure low pressure ELECTRICAL CIRCUIT BATTERY RESISTOR high voltage low voltage
Same shape, different substance. The pump raises pressure so water flows; the battery raises voltage so charge flows. Narrow the pipe and less water gets through — add resistance and less current gets through.
Electrical termUnitWater equivalentWhat it really means
Voltage (V)voltsWater pressureThe push. How hard the source shoves charge around the loop
Current (I)ampsLitres per secondThe flow. How much charge passes a point each second
Resistance (R)ohms (Ω)Pipe narrownessThe obstruction. How hard the material makes it to flow
Power (P)wattsWork the water doesEnergy used per second — light, heat, motion

Where the analogy fails. Water can pour out of an open pipe; electricity cannot. Charge needs a complete loop back to the source, which is why a broken wire stops everything and why a battery with one terminal connected does precisely nothing. Also: a burst pipe wastes water, but a short circuit doesn’t waste electricity — it dumps enormous current through a near-zero resistance and generates enough heat to start a fire.

Ohm’s law: the one equation that matters

Those three quantities aren’t independent. They’re locked together by a relationship discovered by Georg Ohm in the 1820s, and it’s the single most useful equation in electronics.

V = I × R Voltage = Current × Resistance  ·  rearranged: I = V / R  and  R = V / I

In plain language: for a given push, more resistance means less flow. Double the voltage and you double the current; double the resistance and you halve it.

Worked example 1 — how much current through a bulb?

A small bulb with 24Ω of resistance is connected to a 12 V car battery.

I = V / R = 12 / 24 = 0.5 amps

Half an amp flows. Swap in a 6Ω bulb and the current jumps to 2 amps — four times as much, because resistance dropped to a quarter.

Worked example 2 — why thin wires get hot

Power dissipated as heat is P = I² × R. Note the current is squared — doubling the current quadruples the heat.

A thin extension cable might have 0.5Ω of resistance. Run a 10 A heater through it: P = 10² × 0.5 = 50 watts of heat produced in the cable itself. That’s a small heater’s worth of energy, inside a coiled cable, under a rug. This is exactly how extension-lead fires start.

Worked example 3 — what your bulbs actually cost

Power is P = V × I, and your bill is in kilowatt-hours: power × time.

An old 60 W incandescent bulb running 3 hours a day uses 60 × 3 × 365 = 65,700 watt-hours = 65.7 kWh a year. An equivalent 9 W LED uses about 9.9 kWh.

At around 16 cents per kWh, that’s roughly $10.50 a year versus $1.60 — per bulb. Across twenty bulbs in a house, the difference is real money, and none of it required changing a single habit.

What counts as a circuit

A circuit is a complete loop that charge can travel around. Break the loop anywhere and everything stops — which is, incidentally, all a light switch does. It doesn’t “turn on electricity”. It closes a gap in a loop that was otherwise complete.

1 · SOURCE Battery 2 · SWITCH Open = no flow 3 · LOAD Bulb 4 · CONDUCTORS — the wires completing the loop
Every circuit needs these four things: a source of voltage, conductors to carry the charge, a load that does something useful with the energy, and usually a switch to break the loop deliberately.

The load is the point of the whole exercise. It’s where electrical energy converts into something you want — light in a bulb, heat in a kettle, motion in a motor, sound in a speaker. A circuit with no load is a short circuit, and that’s a problem rather than a design.

Series vs parallel: the two ways to wire anything

Once you have more than one component, you face a choice. Chain them one after another, or give each its own branch. Here is what that looks like:

SERIES — one path 12 V 4 V4 V4 V SAME CURRENT EVERYWHERE · VOLTAGE SPLITS PARALLEL — separate paths 12 V 12 V12 V12 V SAME VOLTAGE EACH · CURRENT SPLITS
Three identical bulbs, one 12 V battery, two wiring choices. On the left each bulb gets a third of the voltage and burns dim. On the right each gets the full 12 V and burns bright — but the battery supplies three times the current and drains three times as fast.

That caption contains the entire trade-off, so it’s worth restating: series shares the voltage out, parallel shares the current out. Everything else follows from those two facts.

Series vs parallel, compared properly

tap to switch

Series — components in a single chain

Current: identical through every component. There’s only one path, so whatever leaves the battery must pass through all of them. If 0.5 A flows through the first bulb, 0.5 A flows through the last.

Voltage: divides between components in proportion to their resistance. Three identical bulbs on 12 V get 4 V each.

Resistance: simply adds. Rtotal = R1 + R2 + R3. Three 10Ω bulbs give 30Ω, so less current flows overall.

If one fails: everything stops. The single path is broken and there’s no alternative route.

1 pathroutes for current
Adds uptotal resistance
All offif one fails

Where you actually find it: switches and fuses (deliberately in series so they can break the whole circuit), batteries stacked in a torch to add voltage, LED strips with a current-limiting resistor, and old-style Christmas lights.

Parallel — components on separate branches

Voltage: identical across every branch. Each one connects directly to both sides of the source, so each gets the full 12 V — or the full 120 V or 230 V in your walls.

Current: splits between branches according to what each one draws. A 2 A branch and a 1 A branch means the source supplies 3 A.

Resistance: goes down as you add branches, because you’re adding more routes. 1/Rtotal = 1/R1 + 1/R2. Two 10Ω branches give just 5Ω.

If one fails: the others carry on. The remaining paths are untouched.

Manyroutes for current
Dropstotal resistance
Others fineif one fails

Where you actually find it: every socket and light fitting in your home, car headlights, USB ports on a hub, power strips, and the cells in an electric vehicle battery pack.

Do the maths once and it sticks

Take two 6Ω bulbs and a 12 V battery, and wire them both ways.

QuantityIn seriesIn parallel
Total resistance6 + 6 = 12Ω1/(1/6+1/6) = 3Ω
Current from battery12/12 = 1 A12/3 = 4 A
Voltage across each bulb6 V12 V
Current through each bulb1 A2 A
Power per bulb6 W24 W
ResultBoth dimBoth bright, battery drains 4× faster

Look at the last row. The parallel pair produces eight times the total light output of the series pair — and drains the battery four times as fast. Neither wiring is “better”. They’re different bargains, and which one you want depends entirely on what you’re building.

Series divides the push. Parallel divides the flow. Every other difference is a consequence.

Five real things, and how they’re wired

This is where the theory earns its keep. Each of these is a wiring decision someone made deliberately.

Real-world circuits

tap an example

Christmas lights — the classic series circuit

Old string lights wired fifty small bulbs in series across mains voltage. Each bulb only had to handle a fiftieth of it, which meant cheap, low-voltage bulbs and almost no wiring cost. That was the point.

The price was the failure mode everyone remembers: one filament breaks, the single path is severed, and all fifty go dark with no clue which one failed.

How modern strings fixed it: each bulb contains a tiny shunt — a backup conductor that bridges the gap when the filament burns out, keeping the path intact. The dead bulb goes dark, the rest stay lit. Still a series circuit; just one with a built-in escape hatch.

The catch: with one bulb shunted out, the remaining bulbs share the same voltage between fewer of them, so each runs slightly hotter. Let enough of them fail and the survivors burn out faster and faster.

House wiring — parallel, for four good reasons

Every socket and light in your home sits on its own branch across the same supply. That choice buys four things at once:

  • Full voltage everywhere. Every appliance gets 120 V or 230 V regardless of what else is plugged in.
  • Independence. A blown bulb, or a switched-off lamp, changes nothing elsewhere.
  • Individual control. Each branch can have its own switch.
  • Mixing appliances. A 5 W phone charger and a 2,000 W kettle coexist happily, each drawing what it needs.

The cost: currents add up. On a 120 V circuit rated for 15 A, the ceiling is 1,800 W. A 1,500 W kettle and a 900 W microwave together want 2,400 W — and the breaker trips. That’s not a fault, that’s the breaker doing exactly its job before the wiring overheats.

Car headlights — parallel, for safety

Wired in series, one blown headlight would leave you driving at night with none. In parallel, a failed bulb leaves the other burning at full brightness — degraded, but survivable.

The same reasoning runs through the whole vehicle: indicators, brake lights, wipers, radio all hang in parallel off the 12 V system, each with its own fuse. That fuse, incidentally, is in series with the thing it protects, because a fuse’s job is to break the single path when current runs too high.

The wider pattern: anything safety-critical gets parallel wiring, so that a single failure degrades the system instead of killing it.

A torch — both at once

Stack two AA batteries end to end and you’ve wired them in series: 1.5 V + 1.5 V = 3 V. Series adds voltage, which is why devices needing more push want more cells in a line.

Put two AA cells side by side wired positive-to-positive and you have them in parallel: still 1.5 V, but twice the capacity, so the device runs about twice as long.

Big battery packs use both. An electric vehicle pack might wire cells in series to reach several hundred volts, then wire whole groups of those in parallel for capacity. Series for voltage, parallel for endurance.

Why torch bulbs are in series with the switch: so one switch breaks the only path. Same reason your wall switch works.

Phone chargers and USB — parallel with a limit

Plug three devices into a multi-port charger and each port delivers its own voltage independently — parallel branches. One device finishing its charge doesn’t affect the others.

But the charger has a total current budget. A 65 W charger shared across three ports cannot give all three their full fast-charge current at once, which is why charging slows down as you add devices. The parallel branches are independent in voltage but competing for the same supply.

The same logic scales up: a power strip is parallel sockets sharing one wall circuit. Six free sockets do not mean six appliances’ worth of capacity — the limit lives back at the breaker.

AC and DC, briefly

Everything above works identically for both, but the distinction is worth knowing.

DC

Direct current

Charge flows one way, steadily. Batteries, USB, cars, phones, and everything with a chip in it. Simple, but hard to send over long distances without losses.

AC

Alternating current

Flow reverses direction 50 or 60 times per second. What comes out of your walls, because AC voltage can be stepped up and down with transformers — which is what makes a national grid possible.

WHY AC WON

Transmission

High voltage means low current for the same power, and low current means far less heat lost in the cables. Transformers make that easy for AC and difficult for DC.

THE HYBRID

Your charger

The brick converts AC from the wall into DC for your device. Almost every electronic thing you own is quietly running on DC behind an AC socket.

Electrical safety, with the actual numbers

Here’s the fact that reframes everything: voltage doesn’t hurt you. Current does. Static electricity can hit thousands of volts and merely make you jump, because almost no current flows. A car battery is only 12 V and will not shock you through dry skin — but short its terminals with a spanner and the current can weld metal.

1 mAYou can just feel it
5 mAPainful shock
10–20 mAMuscles lock — you cannot let go
100 mACan stop the heart’s normal rhythm
1–2 ASevere burns, cardiac arrest

Read that middle row again. A tenth of an amp — one thousandth of what a modest household circuit can supply — is potentially fatal. The gap between “you can feel it” and “this could kill you” is about a hundred milliamps, which is nothing at all.

What fuses and breakers actually protect. A 15 A breaker exists to stop your wiring catching fire. It does not protect you — the current that stops a heart is roughly a hundred times below its trip point, so a person can be electrocuted without the breaker noticing anything unusual.

The device that protects people is different: an RCD or GFCI. It compares the current going out with the current coming back, and if even a few milliamps are escaping — through a person, for instance — it cuts power in a fraction of a second. That’s why they’re required near water. If your bathroom or kitchen sockets don’t have one, that’s worth asking an electrician about.

And the obvious one: mains wiring is not a learning project. Batteries, low-voltage kits and USB electronics are genuinely safe to experiment with. Anything connected to your wall is work for a qualified electrician.

Five things people get wrong

“Electricity gets used up as it flows.”

Charge isn’t consumed — exactly as many electrons return to the source as leave it. What gets used up is energy, converted into light, heat or motion by the load. The current arriving at your bulb and the current leaving it are identical.

“Higher voltage always means more dangerous.”

Not on its own. Danger depends on the current that ends up flowing through you, which depends on voltage and the resistance of the path — and skin resistance drops enormously when wet. That’s why a bathroom is more hazardous than a bedroom at exactly the same voltage.

“Adding more bulbs in parallel makes each one dimmer.”

Not from an adequate supply. Each parallel branch still gets full voltage, so each bulb is just as bright — you’re simply drawing more total current. It only dims if the supply can’t keep up, which is why a torch dims as its battery weakens.

“Birds on power lines are safe because of their feet.”

They’re safe because they’re only touching one wire. Current needs a difference in voltage to flow, and both feet are at the same potential. Touch a second wire or a grounded pole at the same time and the outcome is very different — which is why large birds are at more risk than small ones.

“Series is old-fashioned and parallel is modern.”

Neither is obsolete. Series is exactly right for switches, fuses, current-limiting resistors and stacking battery cells for voltage. Parallel is right for independent loads. Real devices use both, often within centimetres of each other.

Check yourself

Five questions. Open each to check — the correct option is marked.

1. Three identical bulbs are wired in series across a 12 V battery. What is the voltage across each?
  • 12 V each
  • 4 V each
  • 36 V total
  • It depends on the switch

Series divides voltage between components. Three identical bulbs share 12 V equally, so 4 V each — which is why they burn noticeably dim.

2. Two 10Ω resistors are wired in parallel. What is the total resistance?
  • 20Ω
  • 10Ω
  • 0.1Ω

Parallel paths reduce total resistance: 1/R = 1/10 + 1/10 = 2/10, so R = 5Ω. More routes means easier flow.

3. Why does a blown bulb in your house not affect the others?
  • Modern bulbs have backup filaments
  • House lights are wired in parallel, so each has its own path
  • The breaker reroutes the current
  • They’re wired in series with a bypass

Each fitting is its own branch across the supply. Break one branch and the rest are untouched.

4. A 24Ω bulb is connected to 12 V. How much current flows?
  • 2 A
  • 0.5 A
  • 288 A
  • 12 A

Ohm’s law: I = V/R = 12/24 = 0.5 A.

5. Which is more dangerous: 20,000 V of static, or 0.1 A through your chest?
  • The static, because the voltage is huge
  • The current — 0.1 A can disrupt the heart
  • Both are equally harmless
  • Neither can harm a person

Static carries a huge voltage but almost no current and no sustained flow. Current through the body is what causes harm, and 100 mA is in the dangerous range.

Frequently asked questions

What is the difference between series and parallel circuits?

A series circuit has one single path, so the same current flows through everything and the voltage divides between components — and if one part fails, everything stops. A parallel circuit has multiple branches, so each gets the full voltage while the current divides, and one failed branch doesn’t affect the others.

Why are houses wired in parallel?

So every socket and light gets the full supply voltage, can be switched independently, and keeps working when something else fails or is unplugged. The trade-off is that currents add up, which is why circuits have breakers.

What is Ohm’s law in simple terms?

Voltage equals current times resistance, or V = I × R. More push gives more flow; more resistance gives less flow. Rearranged as I = V/R, it tells you the current in any simple circuit from just two numbers.

Does adding more devices in parallel drain a battery faster?

Yes. Each branch draws its own current and they add together, so the battery supplies more total current and empties sooner. Each device still runs at full brightness or power until the battery itself starts to sag.

Is it the volts or the amps that kill you?

The current through your body does the damage, but voltage is what drives that current. Roughly 10 to 20 mA is enough to lock your muscles so you cannot let go, and around 100 mA can disrupt the heart. Wet skin drops your resistance dramatically, which is why the same voltage is far more dangerous near water.

The takeaway

Electricity is charge flowing around a complete loop, pushed by voltage and slowed by resistance, with those three quantities tied together by V = I × R. A circuit needs a source, conductors, a load and usually a switch — and break the loop anywhere and it all stops.

The wiring choice is the interesting part. Series puts everything on one path: same current, shared voltage, one failure takes out the lot. Parallel gives everything its own path: same voltage, split current, failures stay local. Your house is parallel so a blown bulb doesn’t darken the kitchen. Your fuse is in series so it can cut everything when it needs to.

Go and look at something nearby with that lens. Why does the switch by the door control only those lights? Why does the fridge keep running when the toaster trips a breaker? Why do two AA cells in a torch make 3 V and not 1.5 V? You now have everything you need to answer all three — and that’s the difference between having read about circuits and actually understanding them.

electricityseries circuitparallel circuitohm’s lawvoltagecurrent

The post How Electricity Works · Circuits Explained (Series vs Parallel) appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/how-electricity-works/feed/ 0 865
The 50/30/20 Budget Rule with Real Numbers https://learnwithexamples.org/50-30-20-budget-rule/ https://learnwithexamples.org/50-30-20-budget-rule/#respond Tue, 04 Aug 2026 18:17:09 +0000 https://learnwithexamples.org/?p=861 Learn With Examples Most budgets fail because they have forty categories and demand perfection. This one has three buckets and tolerates being roughly right. Here is exactly how the maths…

The post The 50/30/20 Budget Rule with Real Numbers appeared first on Learn With Examples.

]]>

Learn With Examples

Most budgets fail because they have forty categories and demand perfection. This one has three buckets and tolerates being roughly right. Here is exactly how the maths lands at five real incomes — including the one where it doesn’t work.

Reading time12 min
LevelBeginner friendly
Includes5 worked incomes

Almost everyone who has tried budgeting has done the same thing: opened a spreadsheet, built thirty-one categories, tracked every coffee for eleven days, then quietly stopped. The failure isn’t discipline. It’s that the system demanded more attention than the problem was worth.

The 50/30/20 rule survives because it asks almost nothing of you. Three buckets. One decision per purchase. And crucially, it doesn’t care whether you spent $40 or $60 on dinner this month, as long as the total for that bucket holds.

The rule comes from All Your Worth, a 2005 book by Elizabeth Warren — then a bankruptcy law professor — and her daughter Amelia Warren Tyagi. It emerged from research into why families went broke, and the answer was rarely lattes. It was fixed commitments quietly growing until nothing flexible remained.

The rule in three lines

Split your take-home pay three ways

50% on needs — the things that cause real trouble if unpaid: housing, food, utilities, transport to work, insurance, minimum debt payments.

30% on wants — everything that makes life enjoyable but wouldn’t end it: eating out, streaming, holidays, hobbies, the nicer phone.

20% on savings and debt repayment — emergency fund, retirement, investments, and any debt payment beyond the minimum.

The percentages apply to take-home pay — what actually lands in your account after tax and deductions. Budgeting your gross salary is the single most common way to get this wrong.

The rule at real incomes

Percentages are easy to nod at and hard to picture. So here are the actual figures, including a realistic list of what the needs half has to absorb.

The rule at five real incomes

pick a take-home

A part-time or entry-level take-home — $2,500 take-home per month.

50%Needs$1,250
30%Wants$750
20%Savings$500

What the needs half has to cover

Rent (shared flat)$700
Groceries$300
Transport pass$120
Phone + internet$70
Basic insurance$60
Still free inside the needs half$0

Putting $500 a month aside for ten years at a 7% average annual return would grow to roughly $86,542, of which $60,000 is your own deposits. Nothing clever — just consistency.

A common single-earner take-home — $3,500 take-home per month.

50%Needs$1,750
30%Wants$1,050
20%Savings$700

What the needs half has to cover

Rent$900
Groceries$350
Car payment + fuel$250
Utilities$130
Insurance$120
Still free inside the needs half$0

Putting $700 a month aside for ten years at a 7% average annual return would grow to roughly $121,159, of which $84,000 is your own deposits. Nothing clever — just consistency.

A mid-career salary after tax — $5,000 take-home per month.

50%Needs$2,500
30%Wants$1,500
20%Savings$1,000

What the needs half has to cover

Mortgage or rent$1,400
Groceries$550
Two cars + fuel$300
Utilities$150
Insurance$100
Still free inside the needs half$0

Putting $1,000 a month aside for ten years at a 7% average annual return would grow to roughly $173,085, of which $120,000 is your own deposits. Nothing clever — just consistency.

A dual-income household with one child — $7,500 take-home per month.

50%Needs$3,750
30%Wants$2,250
20%Savings$1,500

What the needs half has to cover

Mortgage$2,100
Childcare$900
Groceries$800
Cars + fuel$600
Utilities + insurance$500
Over the needs half by$1,150

This is the realistic case the rule struggles with. Childcare alone breaks the 50% line, and no amount of skipping coffee fixes a gap this size. See the section on bending the ratios below.

Putting $1,500 a month aside for ten years at a 7% average annual return would grow to roughly $259,627, of which $180,000 is your own deposits. Nothing clever — just consistency.

A common monthly in-hand salary in India — ₹60,000 take-home per month.

50%Needs₹30,000
30%Wants₹18,000
20%Savings₹12,000

What the needs half has to cover

Rent₹18,000
Groceries₹6,000
Transport / fuel₹3,000
Utilities + phone₹2,000
Insurance premium₹1,000
Still free inside the needs half₹0

Putting ₹12,000 a month aside for ten years at a 10% average annual return would grow to roughly ₹2,458,140, of which ₹1,440,000 is your own deposits.

Two things become obvious once you look at the numbers rather than the percentages.

First, the needs bucket is tight almost everywhere. At $5,000 take-home, half is $2,500 — and in many cities a decent one-bedroom flat plus a car plus insurance is most of that before groceries. The rule isn’t generous; it’s a constraint that mainly binds on housing.

Second, the savings bucket is smaller than it feels and bigger than it looks. $1,000 a month feels enormous when you’re deciding whether to keep a gym membership. Over ten years at a 7% average return, it becomes something in the region of $173,000. The gap between those two feelings is where most financial regret lives.

Need or want? The line that decides everything

This is where the rule gets argued about, and where most people quietly cheat. A useful test: a need is something that causes a real problem within a month if you stop paying it. Rent unpaid becomes eviction. Netflix unpaid becomes mild irritation.

CategoryNeedWantWhere the line sits
HousingRent or mortgageUpgrading for more spaceShelter is a need; a bigger place with the same shelter is a want
FoodGroceriesRestaurants, deliveryFeeding yourself is a need; someone else cooking is a want
TransportGetting to workThe upgraded modelA used hatchback is a need; the leased SUV is partly a want
PhoneA basic planThe premium handsetConnectivity is a need; the newest device is a want
InsuranceHealth, home, vehicleAlmost always a need — it protects everything else
DebtMinimum paymentsMinimums are needs; anything extra counts as savings
ClothingWork-appropriate basicsFashionCovered and presentable is a need; the rest is a want

That last row about debt confuses people, so it’s worth stating plainly: the minimum payment on a loan is a need, and anything you pay above the minimum belongs in the 20% bucket. Paying down a credit card at 22% interest is mathematically identical to earning a guaranteed 22% return. It is saving, just running backwards.

Test yourself on the awkward ones

Your gym membership

Usually a want. Physical activity is a need; a paid membership rarely is, since walking and running cost nothing. The honest exception is if it’s genuinely load-bearing for a health condition or your work. Most people put it under needs because it feels virtuous — that’s the cheat this framework is designed to expose.

Your car

Split it. If you need a vehicle to reach work, the cost of a reasonable, reliable car is a need. The gap between that and a premium lease is a want. A useful method: put the payment on a basic equivalent in needs, and the difference in wants.

Coffee on the way to work

A want, and a small one. Personal finance advice has obsessed over this for years, largely because it’s an easy target. Five dollars a day is around $110 a month — real money, but a rounding error next to a housing decision that’s $400 a month too expensive. Fix the big number first.

Saving for a holiday

A want, not savings. Money you’re setting aside to spend on something enjoyable is deferred spending, not wealth building. Put it in the 30%. The 20% bucket is strictly for things that improve your financial position: emergency fund, investments, extra debt repayment.

Childcare so both parents can work

A need — and often the one that breaks the rule. It’s genuinely required for income to exist, so it belongs in the needs half. In many places it’s simply too large to fit, which is a structural problem, not a budgeting failure.

When the rule doesn’t fit

Any honest treatment of 50/30/20 has to admit where it fails, because for a lot of people it fails immediately.

Expensive cities

Financial guidance often suggests keeping housing near 30% of income. In many major cities the realistic figure is 40 to 50% on its own — leaving the needs half already spent before food.

Low incomes

Below a certain point, needs simply are 70 or 80% of income. Telling someone to save 20% of a wage that barely covers rent isn’t advice, it’s arithmetic denial.

High incomes

The opposite problem. At a high salary, needs may be 25% of take-home — and spending 30% on wants by default is a large amount of money on autopilot.

Irregular income

Freelancers and business owners have no stable monthly figure. The workable fix is to budget percentages against a conservative baseline month, not an average one.

When the ratios need bending

tap a variant

50/30/20 The standard

A balanced starting point when housing genuinely fits inside half your pay.

50%Needs
30%Wants
20%Savings

60/20/20 High cost of living

Rent alone eats most of the needs bucket. Wants shrink first; saving is protected.

60%Needs
20%Wants
20%Savings

70/20/10 Tight, or just starting

Survival first. Ten percent still builds the habit and an emergency buffer.

70%Needs
20%Wants
10%Savings

50/20/30 Catching up or getting ahead

Higher earners, aggressive debt repayment, or saving hard for a deposit.

50%Needs
20%Wants
30%Savings

These are templates, not laws. What matters far more than the exact split is that the third bucket exists and gets funded first.

Notice which bucket moves in each variant. When money is tight, the wants bucket shrinks and savings shrinks last. When money is plentiful, savings expands rather than wants. That priority ordering matters more than the specific numbers, and it’s the part most people invert without noticing.

The percentages are a starting hypothesis. The habit of having a savings bucket at all is the part that changes outcomes.

Setting it up in one evening

StepWhat to doTime
1Find your real monthly take-home — the number that lands in your account, averaged over three months if it varies5 min
2Calculate your three targets: multiply by 0.5, 0.3 and 0.21 min
3Download the last three months of bank and card transactions10 min
4Tag every line as N, W or S. Don’t agonise — first instinct is usually right30 min
5Total each bucket and compare to your targets. This is your actual starting point10 min
6Automate the savings transfer for payday, before you can spend it5 min

Step 6 is doing the heavy lifting, and it’s worth understanding why. Most people treat saving as what’s left at month end — and what’s left is reliably nothing, because spending expands to fill available money. Moving the transfer to payday inverts that: your wants bucket becomes what’s left, and wants are far more elastic than savings.

The three-account setup that makes this automatic. One account receives your salary and pays fixed needs by direct debit. A second gets an automated transfer on payday and is the only account attached to your everyday debit card — that’s your wants money, and when it’s empty, it’s empty. A third holds savings and is deliberately awkward to reach. No tracking app, no discipline required. The structure does the work.

Four mistakes that quietly break it

Budgeting gross pay

Using your headline salary rather than what actually arrives means every bucket is inflated by 20 to 30% and the plan fails in week one.

Forgetting annual costs

Insurance renewals, car servicing, festivals, gifts. Divide the yearly total by twelve and set that aside monthly — otherwise these arrive as “emergencies” that aren’t.

Relabelling wants as needs

The subscriptions, the upgrades, the deliveries. If the needs bucket keeps overflowing, audit what you’ve quietly filed there.

Missing free money

If an employer matches retirement contributions, that match is part of your 20% and it’s an instant guaranteed return. Not claiming it is the most expensive omission on this list.

What to do with the 20%, in order. A common sequence: first, a small starter emergency fund of about one month’s expenses. Second, claim any employer retirement match in full. Third, clear high-interest debt — anything above roughly 8 to 10% beats most expected investment returns. Fourth, build the emergency fund to three to six months. Then invest steadily. Your own circumstances may reorder this, and it’s worth talking through with someone qualified who knows your full picture.

How it compares to other methods

MethodHow it worksEffortBest for
50/30/20Three buckets by percentageLowBeginners, anyone who abandoned detailed budgets
Zero-basedEvery unit of money assigned a job before the month startsHighTight budgets where precision genuinely matters
Envelope / cash stuffingPhysical or digital envelopes per categoryMediumOverspending in specific categories
Pay yourself firstSave a set amount, spend the rest freelyVery lowAnyone who finds any tracking unbearable
80/20Save 20%, don’t track the other 80% at allMinimalPeople whose spending is already under control

The honest comparison: zero-based budgeting produces better results for people who stick with it, and most people don’t. 50/30/20 produces good-enough results with a fraction of the effort, which usually beats a perfect system that gets abandoned in February. If you’ve already failed at a detailed budget twice, that’s information — pick the method that matches the effort you’ll actually sustain.

Check yourself

Five questions. Open each to check — the correct option is marked.

1. The percentages apply to which figure?
  • Gross annual salary
  • Take-home pay after tax and deductions
  • Salary plus expected bonus
  • Household income before rent

Only money that actually reaches your account can be budgeted. Using gross pay inflates every bucket by 20 to 30% and guarantees the plan fails.

2. You pay $200 above the minimum on a credit card. Which bucket?
  • Needs
  • Wants
  • Savings — the minimum is a need, the extra is saving
  • It doesn’t count

Clearing debt at 22% interest is equivalent to a guaranteed 22% return. It’s wealth building, just running in reverse.

3. Take-home pay is $4,200. What is the wants budget?
  • $840
  • $1,260
  • $2,100
  • $420

30% of $4,200. Needs would be $2,100 and savings $840.

4. Rent alone takes 45% of your take-home. What’s the sensible move?
  • Abandon budgeting
  • Count part of the rent as a want
  • Shift to something like 60/20/20 and protect the savings bucket
  • Stop saving until you move

Adjust the ratios to reality rather than pretending. Shrink wants before savings — that priority order is the part that matters.

5. Why automate the savings transfer on payday?
  • Banks pay more interest for it
  • Spending expands to fill whatever is available, so saving last means saving nothing
  • It improves your credit score
  • It’s required by the rule

Saving what’s left over reliably produces nothing left over. Moving the transfer first makes the flexible bucket the leftover instead.

Frequently asked questions

What is the 50/30/20 rule exactly?

A budgeting framework that splits your after-tax income into 50% for needs, 30% for wants, and 20% for savings and debt repayment beyond minimums. It comes from the 2005 book All Your Worth by Elizabeth Warren and Amelia Warren Tyagi.

Is 50/30/20 based on gross or net income?

Net — your take-home pay after tax and deductions. If your employer deducts retirement contributions before you see the money, those already count toward your 20%, so you can treat your target as reduced accordingly.

What if my needs are more than 50% of my income?

Extremely common, especially in high-cost cities. Shift to a variant such as 60/20/20 or 70/20/10 rather than abandoning the framework. Reduce wants before savings, and treat the housing cost itself as the thing to work on over time — it’s the only line big enough to change the picture.

Does the 20% include my employer’s retirement match?

Your own contributions count toward the 20%. The employer’s match is a bonus on top rather than part of your budget, but it’s free money and should be claimed in full before any other savings goal.

Is 50/30/20 actually good, or just popular?

It’s a solid starting framework, not an optimal one. Its strengths are simplicity and sustainability. Its weaknesses are that 20% may be too little for late starters and that the ratios assume housing costs that many people no longer face. Use it as a first structure and adjust from there.

The takeaway

Three buckets, one percentage each, applied to the money that actually reaches your account. Needs are what breaks if unpaid. Wants are everything enjoyable. Savings is anything that improves your financial position, including extra debt repayment.

If the ratios don’t fit your situation, change them — but keep the ordering. Wants flex first, savings flexes last, and the savings transfer leaves your account before you get a chance to spend it. That single mechanical detail does more work than any amount of tracking.

Start with the smallest possible version tonight: find your real take-home figure, multiply it by 0.2, and set up an automatic transfer for that amount on your next payday. Even at half the target, you’ll have built the mechanism — and the mechanism is the part that compounds.

This article is general educational information, not personalised financial advice. Your tax situation, debts, obligations and goals all change what makes sense for you — for decisions of any size, it’s worth speaking with a qualified financial professional who can see your full picture.

50/30/20 rulebudgetingpersonal financesaving moneymoney managementemergency fund

The post The 50/30/20 Budget Rule with Real Numbers appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/50-30-20-budget-rule/feed/ 0 861
Logarithms Explained with Real-Life Examples https://learnwithexamples.org/logarithms-explained/ https://learnwithexamples.org/logarithms-explained/#respond Tue, 04 Aug 2026 17:55:26 +0000 https://learnwithexamples.org/?p=857 Learn With Examples Earthquakes, sound, acidity, star brightness — whenever the world hands us numbers that span a billion-to-one range, we reach for the same trick. Logarithms are that trick,…

The post Logarithms Explained with Real-Life Examples appeared first on Learn With Examples.

]]>

Learn With Examples

Earthquakes, sound, acidity, star brightness — whenever the world hands us numbers that span a billion-to-one range, we reach for the same trick. Logarithms are that trick, and they are far less abstract than school made them look.

Reading time13 min
LevelBeginner friendly
Includes4 live explorers

In March 2011, an earthquake off the coast of Japan registered magnitude 9.1. Four years later, one in Nepal registered 7.8. On paper that looks like a modest difference — one and a bit points on a ten-point scale, the sort of gap you’d shrug at in an exam grade.

It isn’t. The Japanese earthquake released roughly ninety times more energy. Not ninety percent more. Ninety times.

That gap between how the number looks and what it means exists because the magnitude scale is logarithmic. And once you notice one logarithmic scale, you start seeing them everywhere: the decibels on your headphone warning, the pH on a bottle of cleaner, the f-stops on a camera, the octaves on a piano. All the same idea, all for the same reason.

This article explains what a logarithm actually is in plain language, then walks through the three scales people meet most — earthquakes, sound and pH — with real numbers you can check.

The whole idea in one line

A logarithm answers “how many times did I multiply?”

Multiplication asks: I have 10, multiplied by itself 3 times — what do I get? Answer: 1,000.

A logarithm asks the reverse: I have 1,000 — how many 10s did I multiply to get here? Answer: 3. Written log₁₀(1000) = 3.

That’s it. A logarithm is a counter of multiplications, the way division is a counter of subtractions. Everything else follows.

Start with the powers, not the logs

Logarithms feel strange only because they’re usually taught before the thing they undo. So take the multiplication first:

100=1   101=10   102=100   103=1,000   104=10,000 Each step to the right multiplies by ten. The exponent counts the steps.

Now read the same line backwards. Given 10,000, how many steps from 1? Four. So log₁₀(10,000) = 4. Given 100? Two steps, so the log is 2. The logarithm is simply the exponent, extracted and looked at on its own.

Here’s a shortcut that makes base-10 logs concrete forever: for a whole power of ten, the log is the number of zeros. A million has six zeros, so its log is 6. And for anything in between, the log sits between the neighbouring whole numbers — 5,000 is between 1,000 and 10,000, so its log is between 3 and 4 (about 3.7).

100
1log = 0
101
10log = 1
102
100log = 2
103
1,000log = 3
104
10,000log = 4
105
100,000log = 5
106
1,000,000log = 6

Look at the right-hand column. The values explode — 1 to a million — while the logs plod along, 0 to 6. That compression is the entire practical purpose of the tool. A quantity that spans a million-to-one range becomes a scale from 0 to 6, which fits on a chart, on a dial, and in a human head.

Which base? Base 10 is standard for measuring scales, because our number system is decimal. Base 2 appears throughout computing (each step is a doubling — that’s why log₂ shows up in algorithm analysis and in bits of information). Base e ≈ 2.718, the “natural log”, is standard in maths and physics because it makes the calculus clean. Same idea, different step size.

Example 1 — Earthquakes

Earthquake magnitude is the scale most people have heard of and almost nobody reads correctly. The number describes the amplitude of ground motion recorded by instruments, on a base-10 logarithmic scale.

1 step up = 10× the shaking = ~31.6× the energy Energy grows faster than shaking, because it scales with the 1.5 power of magnitude.

So a magnitude 6 shakes the ground ten times as hard as a magnitude 5, and releases about 32 times more energy. Two steps up, from 5 to 7, means 100 times the shaking and about 1,000 times the energy.

What one magnitude step really means

tap a magnitude

Magnitude 4.0 Minor tremor — the kind most seismic regions get weekly

Felt indoors, rattling windows. Thousands occur every year worldwide.

1.0×ground shaking vs M4.0
1.0×energy released vs M4.0
0steps above M4.0

Magnitude 5.5 Comparable to many moderate regional quakes

Furniture moves, weak buildings crack. Locally alarming, rarely deadly.

32×ground shaking vs M4.0
178×energy released vs M4.0
1.5steps above M4.0

Magnitude 6.7 Northridge, California, 1994

Serious damage in built-up areas. The 1994 Northridge earthquake in California measured 6.7 and caused tens of billions of dollars in damage.

501×ground shaking vs M4.0
11,220×energy released vs M4.0
2.7steps above M4.0

Magnitude 7.8 Nepal 2015 · Turkey–Syria 2023

Catastrophic across a wide region. The 2015 Nepal earthquake and the 2023 Turkey–Syria earthquake were both around 7.8.

6,310×ground shaking vs M4.0
501,187×energy released vs M4.0
3.8steps above M4.0

Magnitude 9.1 Tōhoku, Japan, 2011

Among the largest ever recorded. The 2011 Tōhoku earthquake off Japan measured about 9.0–9.1 and triggered the tsunami that struck Fukushima.

125,893×ground shaking vs M4.0
44.7 million×energy released vs M4.0
5.1steps above M4.0

Shaking multiplies by 10 per whole step; energy multiplies by about 31.6, because energy scales with the 1.5 power of magnitude. That is why a 9.1 is not “slightly worse” than a 7.8 — it releases roughly 90 times more energy.

This is why news coverage that says an earthquake was “upgraded from 7.5 to 7.8” is reporting something significant, not a rounding correction. That 0.3 is roughly a tripling of released energy.

It also explains why the scale has no upper limit but rarely exceeds 9.5. The magnitude depends on how much fault surface ruptures, and the planet simply doesn’t contain faults long enough to go much higher. The largest ever instrumentally recorded, in Chile in 1960, was about 9.5.

Example 2 — Sound and decibels

Human hearing is astonishing: the quietest audible sound and the threshold of pain differ in intensity by a factor of about one trillion. Writing that on a volume dial is hopeless. So sound is measured in decibels — a logarithmic scale that turns 1,000,000,000,000 into a tidy 0 to 120.

+10 dB = 10× the sound intensity  ·  +3 dB = the intensity But perceived loudness is different again: roughly +10 dB sounds “twice as loud” to us.
SoundDecibelsIntensity vs a whisperSafe exposure
Rustling leaves20 dB0.1×Unlimited
Whisper, quiet library30 dBUnlimited
Normal conversation60 dB1,000×Unlimited
Busy city traffic85 dB316,000×About 8 hours
Motorbike, food blender94 dB2.5 million×About 1 hour
Rock concert, chainsaw110 dB100 million×Around 2 minutes
Jet engine at 30 m140 dB100 billion×Immediate damage

The third column is the part worth staring at. A rock concert isn’t “a bit louder” than conversation — it delivers roughly a hundred thousand times the sound intensity to your ear. The scale is compressing that so hard it looks almost reasonable.

The safe-exposure column follows from the same maths, and it’s the practical payoff. Hearing-protection guidance commonly halves the permitted exposure time for every 3 dB increase, because 3 dB is a doubling of intensity. That gives roughly 8 hours at 85 dB, 4 hours at 88, 2 hours at 91, 1 hour at 94. A small number on the dial is a big change in your ear.

Why your volume slider feels wrong. Going from 20% to 40% doesn’t sound twice as loud, because perception is itself roughly logarithmic. Doubling the electrical power adds only 3 dB, and 3 dB is barely noticeable. This is also why a 100-watt speaker isn’t twice as loud as a 50-watt one — it’s about 3 dB louder, which most people would describe as “slightly”.

Example 3 — pH and acidity

pH measures how many hydrogen ions are floating in a liquid. Those concentrations vary across an enormous range, so chemists did what physicists did with sound: took the logarithm. With one twist — pH uses the negative log, so that acidic things get small numbers.

pH = −log10[H+]   →   1 step down = 10× more acidic Lower pH means more hydrogen ions. Each whole number is a factor of ten.
SubstancepHHydrogen ions vs pure water
Battery acid010,000,000× more
Lemon juice2100,000× more
Cola2.5~32,000× more
Black coffee5100× more
Pure water7baseline
Human blood7.4~2.5× less
Household bleach131,000,000× less

Now the medical detail that makes this concrete. Human blood is held between about 7.35 and 7.45. Outside roughly 6.8 to 7.8, cells stop functioning and the situation is life-threatening. That sounds like a generous margin until you translate it: the difference between pH 7.4 and pH 6.8 is a four-fold increase in hydrogen ions. Your body defends that range so fiercely precisely because a small pH move is a large chemical one.

The same arithmetic explains why ocean acidification is taken seriously despite unimpressive-looking numbers. Surface ocean pH has fallen from roughly 8.2 to about 8.1 since the industrial era. A tenth of a point — and roughly a 25 to 30 percent increase in hydrogen ion concentration, because 100.1 ≈ 1.26.

On a logarithmic scale, small differences in the number are never small differences in the world. That is the whole reason the scale exists.

The same trick, four more places

Star brightness

Magnitude runs backwards — brighter stars get smaller numbers — and 5 steps equals exactly 100× the brightness, so one step is about 2.512×.

Camera f-stops

Each stop halves or doubles the light. Shutter speeds do the same. Photography is a base-2 logarithmic system with a friendlier name.

Music

An octave doubles the frequency, and every octave is divided into 12 equal ratio steps. Pitch perception is logarithmic, which is why the frets on a guitar get closer together.

Computing

Binary search takes log₂ n steps because it halves the problem each time. A billion items, about 30 steps.

The property that built the modern world

Before calculators, logarithms were not a topic — they were a labour-saving device, and an enormous one. The reason is this identity:

log(a × i) = log(a) + log(i) Multiplication on the inside becomes addition on the outside.

Look up two logs in a table, add them, look the answer back up, and you have multiplied two large numbers without multiplying anything. Navigators, astronomers and engineers used printed log tables and slide rules on exactly this principle for over three centuries, and the Apollo programme was still flying with slide rules in engineers’ pockets.

The same identity is why logarithmic charts are so useful today: on a log axis, anything growing by a constant percentage plots as a straight line. Compound interest, population growth, viral spread — all straight lines on log paper, and their slope tells you the growth rate directly.

The catch with log charts. That same compression can mislead. A curve that is exploding upward looks calm and nearly flat on a logarithmic axis. During the early COVID-19 period, the same case data plotted linearly and logarithmically produced very different emotional reactions from the same facts. Always check which axis you’re reading.

The three rules worth actually remembering

RuleWhat it doesWhere you meet it
log(ab) = log a + log bTurns multiplying into addingSlide rules, combining decibel sources
log(a/b) = log a − log bTurns dividing into subtractingRatios, dB comparisons between two signals
log(an) = n × log aPulls an exponent down into a multiplierSolving for time in growth and decay problems

That third rule is the one that earns its keep in real problems. Suppose an investment grows 7% a year and you want to know when it doubles. You need to solve 1.07t = 2 — and t is stuck up in the exponent where algebra can’t reach it. Take the log of both sides, and the rule pulls it down: t × log(1.07) = log(2), so t = log(2) / log(1.07) ≈ 10.2 years.

Incidentally, that’s where the “rule of 72” comes from — 72 divided by 7 is about 10.3. The mental shortcut is a logarithm in disguise.

Why we bother: three reasons

Compression

A trillion-to-one range becomes 0 to 120. Any scale spanning many orders of magnitude gets a logarithm sooner or later.

Matching perception

Human senses respond to ratios, not differences. Sound, brightness and pitch all work multiplicatively, so a log scale matches how we actually experience them.

Solving for exponents

Whenever the unknown is an exponent — doubling time, half-life, decay — the logarithm is the tool that gets it back down.

Check yourself

Five questions. Open each to check — the correct option is marked.

1. What is log₁₀(100,000)?
  • 4
  • 5
  • 10
  • 100

Count the zeros: five. So you multiply by ten five times to get from 1 to 100,000.

2. How much more energy does a magnitude 7 earthquake release than a magnitude 5?
  • 2 times
  • 100 times
  • About 1,000 times
  • 40 times

Two steps up. Shaking multiplies by 10 twice (100×), and energy by about 31.6 twice — roughly 1,000×.

3. A sound goes from 60 dB to 90 dB. How much has the intensity increased?
  • 1.5 times
  • 30 times
  • 1,000 times
  • 90 times

Every 10 dB is a factor of ten, and that’s three lots of 10 dB — so 10 × 10 × 10. To your ears it sounds roughly eight times louder, because perception compresses it again.

4. Lemon juice is pH 2, black coffee is pH 5. How much more acidic is the lemon juice?
  • 2.5 times
  • 3 times
  • 1,000 times
  • 30 times

Three whole steps down the pH scale, each a factor of ten. Lemon juice has about a thousand times the hydrogen ion concentration.

5. Why does adding logarithms multiply the original numbers?
  • It’s a coincidence of base 10
  • Logs are exponents, and multiplying powers means adding exponents
  • Because logs are always whole numbers
  • It only works for small numbers

10² × 10³ = 10⁵. The exponents add, and a logarithm is the exponent — which is exactly what made slide rules work.

Frequently asked questions

What is a logarithm in simple terms?

It’s the answer to “how many times do I multiply this base by itself to reach that number?” Since 10 × 10 × 10 = 1,000, the base-10 logarithm of 1,000 is 3. It’s the inverse of raising to a power, the way subtraction is the inverse of addition.

Why are earthquake and sound scales logarithmic?

Because the underlying quantities span enormous ranges — sound intensity varies about a trillion-fold across human hearing. A logarithmic scale compresses that into a usable 0 to 120, and it also matches how our senses respond, which is to ratios rather than absolute differences.

What’s the difference between log, ln and log₂?

Only the base. Written plainly, log usually means base 10 and each step multiplies by ten. ln is the natural log, base e ≈ 2.718, standard in calculus and continuous growth. log₂ is base 2, where each step is a doubling — the everyday base in computing.

Where are logarithms used in real life?

Earthquake magnitude, decibels, pH, star magnitude, camera f-stops, musical pitch, the Big O analysis of algorithms, half-life calculations, compound interest, information measured in bits, and every chart with a logarithmic axis.

Can you take the log of zero or a negative number?

No, not with real numbers. No power of 10 produces zero or a negative result — you can get closer and closer to zero (10⁻⁶ is 0.000001) but never reach it. That’s why log curves shoot downward without limit as they approach zero on the left.

The takeaway

A logarithm counts multiplications. That single sentence covers everything above: the magnitude scale counts factors of ten in ground motion, decibels count factors of ten in sound intensity, pH counts factors of ten in hydrogen ions, and an octave counts doublings of frequency.

The practical habit is to translate before reacting. When you see a jump of one on a logarithmic scale, silently multiply by ten instead of adding one. Magnitude 8 versus 7 is not “a bit worse”. pH 5 versus 6 is not “slightly more acidic”. 100 dB versus 90 dB is not “a touch louder”.

Try it on the nearest example to hand: check the pH printed on a bottle in your kitchen, or the decibel figure in your phone’s hearing-safety settings. Convert the number into a multiplication, and you’ll find the reading changes what it means to you. That’s the whole skill.

logarithmsrichter scaledecibelspH scalemathematicsexponents

The post Logarithms Explained with Real-Life Examples appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/logarithms-explained/feed/ 0 857
Correlation vs Causation — Explained with Real Studies https://learnwithexamples.org/correlation-vs-causation/ https://learnwithexamples.org/correlation-vs-causation/#respond Tue, 04 Aug 2026 17:46:49 +0000 https://learnwithexamples.org/?p=853 Learn With Examples Two things moving together is a fact. One thing causing the other is a claim. The gap between them has produced bad headlines, wasted billions in research…

The post Correlation vs Causation — Explained with Real Studies appeared first on Learn With Examples.

]]>

Learn With Examples

Two things moving together is a fact. One thing causing the other is a claim. The gap between them has produced bad headlines, wasted billions in research money, and in a few famous cases, medical advice that actively harmed people.

Reading time13 min
LevelBeginner friendly
Includes5 real studies

In the 1990s, if you were a woman approaching menopause, your doctor may well have suggested hormone replacement therapy — not only for menopausal symptoms, but to protect your heart. The evidence looked strong. Large, careful observational studies following tens of thousands of nurses found that women taking HRT had substantially less coronary heart disease than women who didn’t.

Then, in 2002, a large randomised trial tested it properly. The heart protection wasn’t there. In the group studied, the therapy was associated with increased risks rather than reduced ones, and the trial arm was stopped early.

What went wrong? Nothing, statistically. The correlation was real. Women on HRT genuinely did have fewer heart attacks. But they were also, on average, wealthier, better educated, more likely to exercise, and more likely to see a doctor regularly. The therapy wasn’t protecting their hearts. Their circumstances were — and those same circumstances made them more likely to be prescribed HRT in the first place.

That’s the entire subject of this article, in one real example. Correlation is a measurement. Causation is an explanation. And the distance between them is where most misleading science reporting lives.

The core distinction

Two sentences that are not the same

Correlation: when one thing changes, another tends to change too. It’s an observation about patterns in data. It can be measured, quantified, and it’s often perfectly real.

Causation: changing the first thing makes the second thing change. It’s a claim about mechanism — and it requires far more evidence than a pattern.

Every correlation has at least four possible explanations, and only one of them is “A causes B”. The skill worth building is naming the other three.

First, what correlation actually measures

Correlation is usually expressed as a number between −1 and +1, written r. It tells you two things: the direction of the relationship and how tightly the points cluster around it. It tells you nothing whatsoever about why.

r = 0.0No relationship
r = 0.45Weak / noisy
r = 0.92Strong positive
r = -0.85Strong negative

Two properties matter here. First, a strong correlation is not stronger evidence of causation than a weak one — it’s just a tighter pattern. An r of 0.95 between two unrelated things is still a coincidence, merely a tidier-looking one. Second, correlation as usually measured only detects straight-line relationships. A real, powerful, curved relationship can produce an r near zero.

The number that fooled everyone. In 2012, a paper in a major medical journal noted that a country’s chocolate consumption correlated with its number of Nobel laureates per capita at roughly r = 0.79 — a strikingly strong figure. The author intended it partly as a joke about exactly this problem. It was reported widely and sincerely anyway. National wealth explains both: rich countries buy more chocolate and fund more research.

The five explanations for any correlation

When you see “X is linked to Y”, these are the candidates. Tap through them — each has a real example attached, and getting fluent in the list is most of the skill.

Why might two things move together?

tap an explanation
Pure chance

Coincidence

Compare enough variables and some will line up beautifully by luck alone. This isn’t a rare accident — it’s a mathematical certainty. Test a thousand random pairs at the usual significance threshold and roughly fifty will look “significant” with nothing behind them.

Real example: a 2000 paper in a teaching journal cheerfully demonstrated a statistically significant correlation between stork populations and human birth rates across European countries. The point was satire, and the mechanism is mundane: larger countries have more storks and more babies.

How to spot it: ask whether anyone predicted this relationship before looking. A pattern found by trawling data needs replication in fresh data before it means anything.

Most common culprit

A third variable (confounding)

Something else drives both. This is the explanation behind the majority of misleading health and social science headlines, and it’s the reason the HRT story went wrong.

Real example: a 1999 study reported that infants who slept with a night light were far more likely to become short-sighted. It made headlines worldwide. Follow-up studies the next year found the association vanished once parental short-sightedness was accounted for. Short-sighted parents were both more likely to use night lights and more likely to pass on the genes. The light was innocent.

How to spot it: ask who ends up in each group and why. If the groups differ in ways other than the thing being studied, the third variable is probably doing the work.

Backwards

Reverse causation

The arrow points the other way. B causes A, not A causes B — and in cross-sectional data, both look identical.

Real example: studies repeatedly find that people who skip breakfast weigh more, which produced years of “eat breakfast to lose weight” advice. When researchers actually randomised people to eat or skip breakfast, weight outcomes barely differed. A plausible reading is partly reverse: people already trying to lose weight skip meals.

Another: patients treated in hospital die at higher rates than people who stay home. Hospitals are not the danger — being seriously ill sends you there.

How to spot it: ask which came first. If the data can’t tell you, neither can the headline.

Who got counted

Selection bias

The relationship is an artefact of who ended up in the sample. Nothing is wrong with the analysis; the group being analysed was assembled in a skewed way.

Real example: the long-held belief that moderate drinking protects the heart rests heavily on comparisons with non-drinkers. But the “never drinks” group quietly includes people who stopped because they were already unwell — the “sick quitter” effect. Studies that separate lifelong abstainers from former drinkers, and genetic analyses that sidestep lifestyle entirely, find the protective effect largely disappears.

How to spot it: ask who is missing from the data, and why they’re missing.

Sometimes true

It really is causal

Sometimes A does cause B, and the correlation is the first clue. Dismissing every correlation is as lazy as accepting every one — the job is to work out which is which.

Real example: smoking and lung cancer. When the link emerged in the 1950s, sceptics argued exactly what this article argues: correlation isn’t causation, perhaps some genetic factor caused both the smoking habit and the cancer. That objection was answered not by one study but by a mountain of converging evidence — which is what real causal claims look like.

How to spot it: look for a dose-response pattern, a plausible mechanism, consistency across very different populations, and correct time ordering.

The question is never “is this correlation real?” It usually is. The question is “what else would produce this exact pattern?”

Five real studies, and what actually happened

These aren’t hypotheticals. Each is a case where a genuine correlation led somewhere — sometimes to a correction, once to a Nobel-worthy public health victory.

Case studies

tap a case
Confounding

Hormone therapy and heart disease

What observation suggestedLarge cohort studies through the 1980s and 90s found women on hormone replacement therapy had markedly less coronary heart disease. Guidelines and prescribing practice followed.
What the trial foundThe Women’s Health Initiative randomised trial, reported in 2002, found no cardiac protection. Risks in the studied population went the wrong way and the arm was halted early.

The lesson: the women taking HRT were healthier before they took anything — a pattern researchers now call the healthy-user effect. Careful statistical adjustment did not rescue the finding, because you can only adjust for confounders you thought to measure.

Confounding

Beta-carotene and lung cancer

What observation suggestedPeople eating diets rich in beta-carotene — carrots, leafy greens — had noticeably lower rates of lung cancer. Supplements looked like an obvious public health win.
What the trials foundTwo large randomised trials in the 1990s tested beta-carotene supplements in high-risk smokers. Lung cancer rates went up in the supplement groups, and one trial was stopped early.

The lesson: eating vegetables and swallowing an isolated compound are not the same intervention. The original correlation was probably tracking a whole dietary and lifestyle pattern, not the nutrient. This is one of the clearest cases where acting on a correlation caused measurable harm.

Confounding

Night lights and short-sightedness

What the 1999 study foundChildren who had slept with a night light or room light as infants were several times more likely to be short-sighted later. It was published in a top journal and reported around the world.
What followed in 2000Independent studies failed to replicate the effect once parental short-sightedness was included. Myopic parents were more likely to light the nursery — and more likely to pass on myopia.

The lesson: the fastest way to test a surprising finding is to ask what kind of household does the thing. Habits cluster with genetics, income and education, and any of those can be the real driver.

Selection bias

Moderate drinking and heart health

What decades of data suggestedA J-shaped curve: moderate drinkers appeared to have lower cardiovascular mortality than both heavy drinkers and non-drinkers. It became received wisdom, and a marketing gift.
What better methods suggestSeparating lifelong abstainers from people who quit for health reasons shrinks the benefit substantially. Genetic approaches, which avoid lifestyle confounding by design, find little support for a protective effect.

The lesson: the comparison group matters as much as the exposure group. If your baseline is quietly full of sick people, everything else looks healthy by contrast.

Genuinely causal

Smoking and lung cancer

The correlationStudies from the early 1950s, including a landmark study following British doctors for decades, found heavy smokers had dramatically elevated lung cancer rates.
Why it held upEnormous effect size, a clear dose-response gradient, risk falling after quitting, consistency across countries and study designs, animal and laboratory evidence, and an identified biological mechanism.

The lesson: this is the template for establishing causation without a randomised trial — you cannot ethically assign people to smoke. The criteria used to argue it, published in 1965, are still the standard checklist. Correlation was the starting point, not the conclusion.

Four corrections and one confirmation. That ratio is roughly what the history of nutritional and lifestyle epidemiology looks like — which is a good reason to treat single observational findings as questions rather than answers.

So how is causation established?

Not by one study. By methods that progressively strip away the alternative explanations. Here they are, weakest to strongest.

Anecdote
A story. Zero control over anything.
Cross-sectional
A snapshot. Can’t even establish which came first.
Cohort study
Follows people over time. Fixes time order, not confounding.
Natural experiment
Something outside the researcher’s control split the groups — a policy change, a lottery.
Randomised trial
Assignment by chance, so groups differ only by luck.
Replicated body of evidence
Many designs, many populations, one consistent answer.

Randomisation is the crucial step, and it’s worth understanding why it works. If you flip a coin to decide who gets the treatment, then income, diet, genetics, motivation and every other variable — including ones nobody has thought of — end up distributed roughly evenly between the groups. Statistical adjustment can only handle confounders you measured. Randomisation handles the ones you didn’t.

When you can’t randomise

You cannot assign people to smoke for thirty years, or to grow up poor. Researchers use other tools:

Natural experiments

A policy changes in one state and not its neighbour. Compare the two and the difference approximates a trial nobody ran.

Genetic approaches

Gene variants are allocated at conception, effectively at random. Comparing people by variant sidesteps lifestyle confounding — the method that undercut the drinking story.

Dose-response

If more exposure reliably means more effect, coincidence gets much harder to argue.

Triangulation

Several methods with different weaknesses pointing the same way. Each design’s blind spot is covered by another’s strength.

Reading a headline in thirty seconds

Six questions. You don’t need statistical training for any of them.

Ask thisWhy it mattersWarning sign
Trial or observation?Randomisation is what removes unknown confounders“Linked to”, “associated with”
Who is being compared?Groups may differ in a dozen unmeasured waysSelf-selected groups, volunteers
Could it run backwards?Reverse causation looks identical in a snapshotBoth measured at the same moment
Humans, and how many?Small or animal studies rarely justify adviceUnder a few hundred; mice
Relative or absolute risk?“Doubles your risk” can mean 1 in 100,000 to 2Percentages with no baseline
Has it replicated?A single striking finding is a hypothesis“New study finds”, “first evidence”

A quick vocabulary tell. Careful researchers write “associated with”, “linked to”, or “correlated with” precisely because they mean correlation. Headline writers then translate that into “causes”, “boosts”, “triggers” or “cures”. When the study says one thing and the headline says the other, trust the study — and notice that the gap is usually introduced after the science is finished.

Where this bites outside science

Business decisions

Customers using your new feature retain better — or your most committed customers were always the ones who tried new features.

Marketing spend

Ad spend correlates with sales. It also rises during the seasons when sales rise anyway. A holdout region tells you far more than a chart.

Education

Children in smaller classes score higher — and smaller classes cluster in wealthier districts with more of everything else.

Personal health

You started a supplement and felt better. So did the season, your sleep, and the illness that was ending anyway.

Check yourself

Five questions. Open each to check — the correct option is marked.

1. A study finds people who drink more coffee live longer. What’s the weakest conclusion to draw?
  • Coffee drinkers differ from non-drinkers in other ways
  • Coffee extends lifespan
  • The finding deserves a randomised test
  • Health may affect coffee habits, not the reverse

Observational data supports the association. The causal claim needs the alternatives ruled out first — and sick people often cut down on coffee, which is reverse causation.

2. In the night-light and myopia case, what was the third variable?
  • Screen time
  • Room size
  • Parental short-sightedness
  • The child’s age

Short-sighted parents were both more likely to use night lights and more likely to pass on myopia genetically. Account for the parents and the effect disappears.

3. Why does randomisation remove confounding so effectively?
  • It increases the sample size
  • It balances variables the researchers never even measured
  • It makes the correlation stronger
  • It removes measurement error

Statistical adjustment can only fix confounders you thought to record. Coin-flip assignment spreads the unknown ones evenly too.

4. “Hospital patients die more often than people at home.” What’s the flaw?
  • Coincidence
  • Reverse causation — illness sends people to hospital
  • The sample is too small
  • There’s no flaw

The arrow runs backwards. Being seriously ill causes hospital admission; the hospital didn’t cause the illness.

5. Why was smoking accepted as causing lung cancer without a randomised trial?
  • The correlation was very strong
  • A single large study settled it
  • Many independent lines of evidence converged — dose-response, mechanism, consistency, risk falling after quitting
  • Governments decided it

Strength alone was never sufficient — sceptics raised exactly the confounding objection. It was the convergence of many designs that closed the case.

Frequently asked questions

What is the difference between correlation and causation in simple terms?

Correlation means two things move together in the data. Causation means changing one actually produces a change in the other. Correlation is measured; causation is inferred, and it needs the alternative explanations ruled out first.

Can correlation ever prove causation?

Not on its own. But correlation plus dose-response, correct time ordering, a plausible mechanism, consistency across different populations and study designs, and preferably experimental evidence together can establish causation convincingly — as happened with smoking.

What is a confounding variable?

A third factor that influences both things being studied, creating a correlation between them that isn’t causal. Wealth confounds chocolate consumption and Nobel prizes; parental short-sightedness confounded night lights and myopia.

Why are randomised controlled trials considered the gold standard?

Because assignment by chance makes the groups comparable on everything, including factors nobody thought to measure. That’s the one thing statistical adjustment of observational data can never fully achieve.

Does “correlation is not causation” mean I should ignore observational studies?

No. Observational research is how most hypotheses start, and for many questions it’s the only ethical option. Treat a single observational finding as a good question rather than a settled answer, and give much more weight to results that replicate across different methods.

The takeaway

Correlation is a fact about data. Causation is a claim about the world. Moving from one to the other requires ruling out coincidence, third variables, reverse causation and selection bias — and the historical record shows how often that fails, even in careful hands, even in top journals.

The HRT story is worth keeping in mind precisely because nobody was careless. Good researchers, huge samples, real statistical rigour, and the answer was still wrong — because the women who took the therapy were different from the women who didn’t in ways the data never captured.

You don’t need to become a statistician. You need one reflex: when you read that X is linked to Y, pause and ask what else could produce that exact pattern. Usually you’ll think of something within about ten seconds. That reflex is worth more than any formula in this article.

correlation vs causationstatisticsconfoundingresearch methodscritical thinkingdata literacy

The post Correlation vs Causation — Explained with Real Studies appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/correlation-vs-causation/feed/ 0 853
Ransomware: How It Works and How It Spreads https://learnwithexamples.org/ransomware-how-it-works-spreads/ https://learnwithexamples.org/ransomware-how-it-works-spreads/#respond Tue, 04 Aug 2026 17:40:10 +0000 https://learnwithexamples.org/?p=848 Learn With Examples It almost never starts with a hacker in a hoodie breaking down a digital door. It starts with one password, one unpatched server, or one convincing email…

The post Ransomware: How It Works and How It Spreads appeared first on Learn With Examples.

]]>

Learn With Examples

It almost never starts with a hacker in a hoodie breaking down a digital door. It starts with one password, one unpatched server, or one convincing email — and by the time the ransom note appears, the attackers have often been inside for weeks.

Reading time13 min
LevelBeginner friendly
Includes5 real case files

Picture a Friday afternoon at a mid-sized logistics company. An accounts clerk gets an email that looks like an unpaid invoice from a supplier she deals with weekly. She opens the attachment. Nothing visibly happens, so she carries on with her day and goes home.

Nineteen days later, at 2:14 on a Sunday morning, every file server in the company locks simultaneously. Spreadsheets, delivery schedules, the customer database, even the backups on the network drive — all replaced with scrambled files and a text document explaining where to send the Bitcoin.

Those nineteen days are the part almost nobody pictures. Ransomware is not a bomb that goes off when you click. It is a burglary that ends with the locks being changed — and the burglary itself takes days or weeks, during which the intruders are quietly reading your email, finding your backups, and stealing your data before a single file is ever encrypted.

This article explains what actually happens in those weeks, how the infection gets in, and what five of the most damaging real-world attacks teach us. Everything here is defensive knowledge: how the attack chain works, so you can recognise and break it.

In one paragraph

What ransomware is

Ransomware is malicious software that makes your data unusable and sells it back to you. It scrambles your files with strong encryption, and the key needed to unscramble them is held by the attacker until you pay. Modern attacks add a second lever: before encrypting anything, they copy your sensitive data out and threaten to publish it. So even a company with perfect backups still faces a leak.

The important part for defenders: encryption is the last step. Everything before it — the break-in, the spread, the theft — is where an attack can still be stopped.

Why it stopped being about locked files

Early ransomware was crude and indiscriminate: infect as many home computers as possible, demand a few hundred dollars each, hope enough people pay. Then criminals noticed that a hospital or a factory will pay vastly more than a thousand individuals, and the whole model changed.

  • 2013 — 2015 · SPRAY AND PRAYMass email campaigns hitting individuals for a few hundred dollars. Good backups were a complete defence.
  • 2016 — 2018 · BIG GAME HUNTINGAttackers begin targeting organisations deliberately, studying revenue and pricing the ransom to fit. Demands jump into six and seven figures.
  • 2019 · DOUBLE EXTORTIONThe pivotal change: steal the data first, then encrypt. Now backups alone don’t save you — refusing to pay means your customers’ data gets published.
  • 2020 onwards · RANSOMWARE AS A SERVICEThe criminal ecosystem specialises. One group writes the software, others break into networks, others negotiate. You no longer need technical skill to run an attack.
  • 2023 onwards · EXTORTION WITHOUT ENCRYPTIONSome groups skip the encryption entirely. Stealing the data and threatening to leak it is quieter, faster, and often just as profitable.

That progression matters because it changes what “being prepared” means. Backups answer the 2015 problem. They do nothing about the 2019 one. A company today needs to survive both losing access to its data and having that data published.

The attack chain, stage by stage

Nearly every serious ransomware incident follows the same six stages. Tap through them — each stage is a place where the attack could still be caught, and knowing that is the entire point of learning the sequence.

The six stages

tap a stage
Stage 1 · Initial access

Getting through the front door

Three routes account for the overwhelming majority of break-ins: a stolen or guessed password on a remote access system without multi-factor authentication, an unpatched internet-facing server, and a convincing email that persuades someone to open something.

Real example: the Colonial Pipeline attackers in 2021 didn’t need anything sophisticated. They used a password for a disused VPN account that appeared in a batch of leaked credentials. That account had no multi-factor authentication on it.

Where it breaks: MFA everywhere, prompt patching of anything reachable from the internet, and shutting down accounts nobody uses.

Stage 2 · Establishing a foothold

Making sure they can get back in

The first thing an intruder does is make their access durable. They set up a second and third way back into the network, so losing one doesn’t lock them out. Increasingly they use ordinary administration tools that already exist on your systems, precisely because those tools don’t look suspicious.

Why it matters: this is why “we changed that password” is not a fix. If someone has been inside for two weeks, the door they came through is rarely the only one still open.

Where it breaks: endpoint monitoring that flags unusual use of legitimate admin tools, and alerting on new accounts or scheduled tasks appearing.

Stage 3 · Reconnaissance

Reading your organisation

Now they explore. Where are the file servers? Which systems would hurt most if they stopped? Where are the backups, and are they connected to the network? They read finance documents to work out what you can afford, and they check whether you have cyber insurance — because that tells them your ceiling.

The uncomfortable statistic: attackers commonly spend days to weeks inside a network before triggering anything. During that window there is nothing to see unless you are looking.

Where it breaks: network segmentation, so that compromising one laptop doesn’t grant a tour of everything, and monitoring for unusual internal scanning.

Stage 4 · Privilege escalation

Becoming the administrator

One ordinary user account isn’t enough to encrypt a company. The goal is domain administrator — the keys to everything. Attackers get there by harvesting credentials left in memory on shared machines, exploiting misconfigured permissions, or simply finding a password in a document called something like server-passwords.xlsx.

Real example: in the 2023 attack on MGM Resorts, attackers reportedly phoned the IT help desk, impersonated an employee, and talked their way into a credential reset. No exploit required — just a convincing phone call.

Where it breaks: least privilege, separate accounts for admin work, and strict identity verification for help-desk resets.

Stage 5 · Data theft

Copying everything out first

Before anything is locked, the valuable data leaves — customer records, contracts, HR files, source code. Often it is uploaded to ordinary cloud storage services, because that traffic blends in with normal business use.

Why this stage exists: it is the insurance policy against your backups. If you restore everything and refuse to pay, they still hold your data and can publish it. This is what turned ransomware from an IT problem into a legal and reputational one.

Where it breaks: monitoring for large outbound transfers, and controls on which cloud services can be reached from inside the network.

Stage 6 · Encryption and extortion

The part everybody pictures

Only now do files get encrypted, usually at night or over a holiday weekend when nobody is watching. Backups are deliberately destroyed or encrypted first. Then the ransom note appears, typically with a countdown and a link to a private chat with the attackers.

The timing is deliberate: attacks cluster around Friday nights, public holidays, and long weekends. Fewer staff, slower detection, more hours to finish the job.

Where it breaks: offline or immutable backups that cannot be reached from the network, and a tested recovery plan — the emphasis being firmly on tested.

Stages 1 to 5 are invisible to most organisations. Stage 6 is impossible to miss. Nearly all the opportunity to prevent a disaster sits in the part nobody sees.

By the time you can see ransomware, you are not preventing an attack any more. You are recovering from one that finished.

How it actually spreads

Two different questions get muddled here: how attackers get into an organisation, and how the infection moves within one. They have different answers.

Getting in: the three doors

Where attacks begin

industry incident reports, broadly consistent year to year
Stolen credentials & exposed remote accesslargest single category
Unpatched internet-facing softwarevery large
Phishing & malicious attachmentsvery large
Supply chain & trusted software updatesrarer, far bigger blast radius

Proportions vary between reports and years, but the ranking is stubbornly stable: credentials and unpatched internet-facing systems consistently beat phishing as the way in.

That last point surprises people, because security training focuses so heavily on suspicious emails. Phishing is a genuine and major route — but a remote access portal with no multi-factor authentication is a permanently open door that requires no one to make a mistake at all. It just needs one password to appear in any of the billions already circulating from previous breaches.

Moving inside: how one machine becomes five hundred

METHOD 01

Stolen credentials

The most common by far. With an administrator password, attackers simply log in to other machines. Nothing is “hacked” — they use the same tools your IT team does.

METHOD 02

Management tools

Software built to push updates to every computer at once is perfect for pushing ransomware to every computer at once. Attackers actively hunt for it.

METHOD 03

Network shares

Mapped drives are a gift. Anything a compromised user can write to, ransomware can encrypt — including that “backup” folder on the shared server.

METHOD 04

Worm-like exploits

Rare but devastating. A flaw needing no human interaction lets the infection jump machine to machine automatically, in minutes rather than days.

The backup trap. A backup drive that is permanently connected and writable is not a backup as far as ransomware is concerned — it is just more files to encrypt. Attackers look for backups deliberately and destroy them first, because a company that can restore is a company that won’t pay. Backups must be offline, immutable, or in a separate account the compromised network cannot reach.

Five real attacks, and what each one teaches

These are the incidents worth knowing, not because they were the most sophisticated, but because each exposed a different weakness.

Case files

tap a case
May 2017 · global

WannaCry — the patch nobody applied

WannaCry spread itself. It used a flaw in an old Windows file-sharing protocol that let it jump from machine to machine with no human involvement, infecting well over 200,000 computers across 150 countries in a couple of days. Britain’s NHS was hit hard: ambulances diverted, thousands of appointments and operations cancelled.

The detail that stings: Microsoft had released the patch two months earlier. Every organisation that applied updates promptly was immune. The outbreak was eventually slowed when a researcher registered a domain name found in the code that acted as an accidental kill switch.

Lesson: patching internet-facing and networked systems is not administrative housekeeping. It is the single highest-return security activity most organisations can do.

June 2017 · global

NotPetya — the one that wasn’t really ransomware

NotPetya arrived through a poisoned update to a Ukrainian tax accounting package that most companies doing business in Ukraine were required to use. It looked like ransomware and demanded payment — but the encryption was designed to be irreversible. Paying achieved nothing. It was destruction dressed as extortion.

Shipping giant Maersk lost virtually its entire global IT estate and rebuilt roughly 4,000 servers and 45,000 computers in about ten days. Total worldwide damages have been estimated at around $10 billion, making it the costliest cyberattack in history.

Lesson: software you trust and update automatically is a route into your network. And a “ransom” is not always a real offer — recovery capability matters more than negotiating ability.

May 2021 · United States

Colonial Pipeline — one password, one country queuing for fuel

The largest fuel pipeline in the US, carrying nearly half the East Coast’s supply, shut down for around six days. Panic buying emptied filling stations across several states and a state of emergency was declared.

The entry point was a single VPN account that was no longer in use, protected by a password that had leaked in an earlier breach elsewhere, with no multi-factor authentication. The company paid roughly $4.4 million in Bitcoin; US authorities later recovered a substantial portion of it. Notably, the pipeline itself was shut down as a precaution — the billing systems were the ones affected.

Lesson: disused accounts are live doors, reused passwords eventually surface in a leak, and MFA on remote access would have ended this attack before it began.

July 2021 · global

Kaseya — attacking a thousand companies at once

Rather than break into companies one by one, attackers compromised a tool that IT service providers use to manage their clients’ computers remotely. The ransomware was then pushed out through that trusted channel to roughly 1,500 downstream businesses in a single stroke — including a Swedish supermarket chain that had to close around 800 stores because its tills stopped working.

The attack landed on the Friday of a long holiday weekend in the US, when response teams were thinnest. That timing was not a coincidence.

Lesson: your security includes your suppliers’ security. Any vendor with remote access into your systems is part of your attack surface.

February 2024 · United States

Change Healthcare — the most expensive lesson in MFA

Change Healthcare processes a large share of American medical claims. When it was hit, pharmacies could not verify insurance, providers could not bill, and parts of the healthcare payment system stalled for weeks. Parent company UnitedHealth ultimately reported costs running into billions of dollars, and the breach affected the data of well over a hundred million people.

The way in was a remote access portal without multi-factor authentication. A $22 million ransom was paid for a promise to delete the stolen data — then the criminal group collapsed in an apparent exit scam, the stolen data moved to a second gang, and a fresh extortion demand followed.

Lesson: paying buys a promise from criminals, not a guarantee. And when one company sits at the centre of an industry, its security is everyone’s problem.

Why it keeps happening: the business behind it

Ransomware persists because it has become an industry with specialised roles, not a hobby for lone hackers.

RoleWhat they doWhy it matters
DevelopersBuild and maintain the software, run the leak site and payment infrastructureThey rent it out and take a cut — they rarely attack anyone directly
AffiliatesDo the actual break-ins using the rented toolkitYou don’t need technical depth to run an attack any more
Access brokersBreak into networks and sell that access to whoever wants itYour network can be sold as a product before anyone encrypts anything
NegotiatorsHandle the victim chat, apply pressure, arrange paymentProfessionalised, scripted, and experienced at reading desperation

This division of labour explains a lot. It explains why attacks are so consistent in method — affiliates follow playbooks. It explains why groups reappear under new names after being disrupted — the people move, the brand changes. And it explains why paying keeps the machine running: every payment funds the next round of development.

Should victims pay?

This is a decision for the organisation with legal counsel and law enforcement, and the honest position is that there is no clean answer. What can be said factually:

  • Paying buys a decryption key that often works imperfectly — restoration from decryptors is typically slow and incomplete.
  • Paying for deletion of stolen data buys a promise. The Change Healthcare case shows what that promise can be worth.
  • Organisations that pay are disproportionately targeted again, sometimes by the same affiliate.
  • In some jurisdictions, payments to sanctioned groups carry legal exposure of their own.
  • Recovery from clean backups is usually faster than decryption — when those backups exist and have been tested.

Which is why the real decision is made long before the ransom note: an organisation that can restore has options, and an organisation that cannot has only one.

Defending against it, in priority order

Tick these off as you read — they’re ordered by how much protection they buy per unit of effort, and the first three stop the majority of real-world attacks outright.

Defence checklist

tap to tick

If it happens anyway. Isolate affected machines from the network but don’t power them off — useful evidence lives in memory. Don’t delete anything. Contact law enforcement and your cyber insurer early, since both may have resources and requirements you need. Assume data was stolen and plan your notifications accordingly. And resist the urge to restore in a rush: rebuilding into a network the attacker still has access to is a well-trodden way to be encrypted twice.

Four myths worth deleting

“We’re too small to be a target.”

Most attacks aren’t targeted at all — they’re opportunistic, driven by automated scanning for exposed systems and weak credentials. Small organisations are attacked constantly, precisely because they’re less likely to have monitoring, and a smaller ransom is still profit for a group running dozens of attacks a month.

“Our antivirus will catch it.”

Antivirus catches known malicious files. It struggles with an attacker who logs in with valid credentials and uses your own administration tools — which is exactly what most modern intrusions look like. Detection has to cover behaviour, not just files.

“We have backups, so we’re fine.”

Backups solve half the problem, and only if they’re offline and tested. They do nothing about stolen data being published, and connected backups are routinely destroyed by the attacker before encryption begins.

“It happens instantly when someone clicks.”

Encryption is the final act, often days or weeks after the initial break-in. That gap is the defender’s real opportunity — every attack that gets detected during it is an attack you never hear about.

Check yourself

Five questions. Open each to check — the correct option is marked.

1. At what point in a ransomware attack does encryption happen?
  • Immediately when someone clicks a link
  • Last, often days or weeks after the initial break-in
  • Before the attacker gets administrator access
  • Only after the ransom is refused

Encryption is stage six of six. Everything before it is invisible to most organisations, and it’s where the attack can still be stopped.

2. What was the entry point in the Colonial Pipeline attack?
  • A zero-day exploit in industrial control systems
  • A malicious email attachment opened by an executive
  • A disused VPN account with a leaked password and no MFA
  • A compromised software update

No sophistication required. One old account, one reused password, no second factor — and half the US East Coast’s fuel supply stopped moving.

3. Why do attackers steal data before encrypting it?
  • Encryption doesn’t work without a copy
  • So that good backups don’t remove the victim’s reason to pay
  • To test whether the files are valuable
  • It makes the encryption faster

Double extortion is the answer to backups. Restore everything and refuse to pay, and they publish your customers’ data instead.

4. Which single control would have prevented the most cases in this article?
  • Antivirus software
  • Employee phishing training
  • Multi-factor authentication on internet-facing access
  • A stronger firewall

Colonial Pipeline and Change Healthcare both came down to remote access without MFA. It’s unglamorous and it’s the highest-value control most organisations can deploy.

5. Why is a permanently connected backup drive a problem?
  • It wears out faster
  • Ransomware encrypts anything writable, and attackers destroy backups first
  • It slows the network down
  • It isn’t a problem

Attackers hunt for backups deliberately, because a company that can restore is a company that won’t pay. Offline or immutable is the requirement.

Frequently asked questions

How does ransomware get onto a computer?

Most commonly through stolen or guessed credentials on remote access systems without multi-factor authentication, through unpatched internet-facing software, or through a convincing email attachment or link. Less often, through a compromised software update from a trusted supplier.

How long does a ransomware attack take?

The encryption itself takes hours. The intrusion leading up to it commonly takes days to weeks, during which attackers explore the network, escalate their access, and copy data out. The visible part is the very end of a long, quiet process.

Can encrypted files be recovered without paying?

Sometimes. Free decryptors exist for older or flawed ransomware families, and projects like No More Ransom collect them. For current strains with correctly implemented encryption, breaking it is not realistic — recovery means restoring from backups.

Does paying the ransom actually work?

Partially, often. Decryption keys usually work but restoration is typically slow and imperfect, and promises to delete stolen data are unverifiable — the Change Healthcare case ended with a second gang demanding a second ransom for the same data. Paying also marks an organisation as willing to pay.

What should a small business do first?

Turn on multi-factor authentication everywhere, especially email and any remote access. Then get backups that are disconnected from the network and actually test restoring from them. Those two steps cost very little and remove the majority of realistic attack paths.

The takeaway

Ransomware is not really a story about encryption. It’s a story about access — how someone gets in, how far they can move once inside, and how long they can stay before anyone notices. The encryption at the end is just the invoice.

Which is why the defences that matter are so unglamorous. Not clever software, but multi-factor authentication on the front door, patches applied on time, backups the attacker can’t reach, and a plan that has been rehearsed before it’s needed. Colonial Pipeline, Change Healthcare, WannaCry — every one of them turns on something a checklist would have caught.

If you take one action after reading this, make it the smallest one: check whether multi-factor authentication is switched on for your email and any remote access you use. That single setting appears in the story of more prevented attacks than anything else in this article.

ransomwarecybersecuritymalwarephishingdata breachMFA

The post Ransomware: How It Works and How It Spreads appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/ransomware-how-it-works-spreads/feed/ 0 848
Big O Notation Explained with Real Examples https://learnwithexamples.org/big-o-notation-explained/ https://learnwithexamples.org/big-o-notation-explained/#respond Tue, 04 Aug 2026 17:20:21 +0000 https://learnwithexamples.org/?p=842 · Learn With Examples · Big O is not maths for its own sake. It is a one-line answer to the only question that matters when your data grows: does…

The post Big O Notation Explained with Real Examples appeared first on Learn With Examples.

]]>

· Learn With Examples ·

Big O is not maths for its own sake. It is a one-line answer to the only question that matters when your data grows: does this code get a little slower, or does it fall off a cliff?

Reading time13 min
LevelBeginner → Practical
Includes3 explorers

Here’s the moment Big O suddenly makes sense to people. I once watched a junior developer ship a feature that checked a list of orders for duplicates. It worked perfectly. Tests passed, code review passed, demo went beautifully. Three months later the same feature took nine minutes to load and brought a support queue to a standstill.

Nothing had changed in the code. The only thing that changed was the number of orders — from about 300 to about 40,000. His approach compared every order against every other order, so the work grew with the square of the list. Going 130 times bigger made it roughly 17,000 times slower.

Big O notation is how you spot that in advance, in about ten seconds, without running anything. It’s a shorthand for describing how the work an algorithm does grows as its input grows — and once you can read it, that nine-minute disaster becomes something you catch during code review instead of during an outage.

The one-paragraph version

What Big O actually says

Big O describes the growth rate of an algorithm’s work as the input size n increases. It deliberately ignores hardware, programming language, and constant factors, because those change by a factor of two or ten — while growth rates change by a factor of millions.

O(1) means the input size doesn’t matter at all. O(log n) means work barely grows. O(n) means it grows in step. O(n²) means doubling the input quadruples the work. That last one is where systems die.

The idea in one picture

Everything else in this article is detail attached to this one shape. Watch what happens to each line as you move right along the horizontal axis — that axis is your data growing over the life of your product.

input size n → operations → O(n log n) O(n²) O(n) O(log n) O(1)
The same five classes, drawn to scale. Notice that near the left edge they’re all bunched together — which is exactly why a slow algorithm looks fine in testing and only reveals itself in production.

That bunching on the left is the trap. With 20 test records, an O(n²) function and an O(n) function both return instantly. There is no observable difference. The difference only exists at scale, which means you cannot discover it by testing on small data — you have to reason about it. That reasoning is what Big O gives you.

The classes, with examples you’ll recognise

There are only six you meet regularly. Tap through them — each one has a plain-English meaning, a real piece of code, and an everyday analogy that makes the growth rate obvious.

Complexity explorer

tap a class

O(1)  Constant time

The work never changes, no matter how big the input gets. Looking up one item takes the same time in a list of ten or ten million.

Real example: fetching a value from a dictionary or hash map, reading array[500], pushing onto a stack, checking whether a number is even.

Analogy: opening a specific page in a book when you already know the page number. The book’s thickness is irrelevant.

O(log n)  Logarithmic time

Each step throws away half the remaining work. Doubling the input adds just one extra step — which is why this class stays fast at absurd scales.

Real example: binary search in a sorted array, finding a record in a balanced tree index, the depth of a database B-tree.

Analogy: the guessing game. “Is it higher or lower?” finds a number between 1 and a million in 20 guesses, because each guess halves the range.

Bars scaled to their own maximum — notice how the growth flattens out almost immediately.

O(n)  Linear time

Work grows in lockstep with the input. Twice the data, twice the time. Perfectly respectable, and often unavoidable — if you must look at every item, you cannot beat linear.

Real example: finding the largest number in an unsorted list, counting words in a document, validating every row of a CSV.

Analogy: reading every name on a guest list to find one person. A list twice as long takes twice as long.

O(n log n)  Linearithmic time

A linear pass repeated across a logarithmic number of levels. Slightly worse than linear, dramatically better than quadratic, and the practical ceiling for comparison-based sorting.

Real example: merge sort, quicksort on average, and every sort() in every standard library you have ever used.

Analogy: splitting a deck of cards in half repeatedly, then merging the piles back in order. You touch every card at each of about ten levels.

O(n²)  Quadratic time

Every item is compared against every other item. Double the input and the work quadruples. This is the class that quietly kills features six months after launch.

Real example: the duplicate-checking loop from the opening story, bubble sort, comparing every pair of records to find matches, a nested loop over the same list.

Analogy: every guest at a party shaking hands with every other guest. Ten guests means 45 handshakes; a hundred guests means 4,950.

O(2ⁿ)  Exponential time

Adding a single item doubles the work. Usable for tiny inputs and hopeless past roughly 40 items — at n = 60 you are past the age of the universe.

Real example: generating every possible subset, naive recursive Fibonacci, brute-forcing a password, solving the travelling salesman by trying every route.

Analogy: the grain-of-rice-on-a-chessboard story. Doubling per square sounds harmless until square 64 needs more rice than exists on Earth.

Why the difference is so violent

Abstract symbols don’t frighten anyone. Actual step counts do. Pick an input size below and look at how the same six classes behave on it.

Steps required, by input size

pick an n
ClassTypical operationSteps for n = 10Time at 1 billion ops/sec
O(1)Look up a key in a hash map1instant
O(log n)Binary search a sorted list3instant
O(n)Scan every item once10instant
O(n log n)A good sort33instant
O(n²)Compare every pair100instant
O(2ⁿ)Try every subset1,0241 microseconds
ClassTypical operationSteps for n = 100Time at 1 billion ops/sec
O(1)Look up a key in a hash map1instant
O(log n)Binary search a sorted list7instant
O(n)Scan every item once100instant
O(n log n)A good sort664instant
O(n²)Compare every pair10,00010 microseconds
O(2ⁿ)Try every subsetbeyond countinglonger than the universe has existed
ClassTypical operationSteps for n = 1,000Time at 1 billion ops/sec
O(1)Look up a key in a hash map1instant
O(log n)Binary search a sorted list10instant
O(n)Scan every item once1,0001 microseconds
O(n log n)A good sort9,96610 microseconds
O(n²)Compare every pair1.0 million1 ms
O(2ⁿ)Try every subsetbeyond countinglonger than the universe has existed
ClassTypical operationSteps for n = 1,000,000Time at 1 billion ops/sec
O(1)Look up a key in a hash map1instant
O(log n)Binary search a sorted list20instant
O(n)Scan every item once1.0 million1 ms
O(n log n)A good sort19.9 million20 ms
O(n²)Compare every pair1 trillion17 minutes
O(2ⁿ)Try every subsetbeyond countinglonger than the universe has existed

Swipe the table sideways on a narrow screen →

One billion operations per second is roughly a fast modern CPU doing simple work. The point is not the exact seconds — it is how violently the bottom two rows change as n grows.

Read across the O(n²) row as you switch sizes. At a thousand items it’s a million steps — a blink. At a million items it’s a trillion steps, and your feature is now a fifteen-minute job that times out. Nothing about the code changed. Only n did.

A faster computer buys you a constant factor. A better algorithm buys you a different curve. Only one of those scales.

The two rules that make Big O readable

Big O deliberately throws information away, and the two things it throws away trip up every beginner. Both rules exist for the same reason: at large n, only the fastest-growing part matters.

Rule 1 — drop the constants

An algorithm that does 3n + 12 operations is written O(n), not O(3n + 12). That looks like cheating, but consider: at a million items, 3n is three million and is a trillion. The constant 3 is noise beside that gap. And constants change anyway — a faster CPU, a better compiler, or a different language shifts them freely, while the shape of the curve does not.

Rule 2 — keep only the dominant term

An algorithm costing n² + 500n + 9000 is O(n²). At n = 10, the 500n part is actually bigger. At n = 10,000, the term is 100 million and the rest is five million — the square has swallowed everything. Big O describes where things end up, not where they start.

The honest caveat. Those discarded constants sometimes matter enormously in practice. An O(n log n) algorithm with a huge constant can lose to an O(n²) one on lists of thirty items — which is exactly why real sorting libraries switch to insertion sort for small chunks. Big O tells you what happens as data grows. It does not promise which is faster today, on your data, at your size.

How to find the Big O of code you’re looking at

You don’t need to count operations. Four patterns cover almost everything you’ll meet in ordinary application code.

What you seeWhat it meansResult
Statements one after anotherCosts add, then the biggest one winsO(a) + O(b) → the larger of the two
A loop over the inputBody runs n timesn × cost of the body
A loop inside a loopCosts multiplyUsually O(n²)
Input halves each stepOnly log₂n steps possibleO(log n)

Two things people get wrong constantly. First, two loops side by side are not quadratic — they’re n + n = 2n, which is linear. Nesting is what multiplies, not adjacency. Second, a loop with a fixed bound isn’t linear: looping 100 times regardless of input is constant work, because 100 doesn’t grow with n.

Five snippets — guess before you open

Read each one, decide, then open it. These are the shapes that actually appear in real code.

1. Checking a list for duplicates the obvious way
for i in range(len(orders)): for j in range(i + 1, len(orders)): if orders[i].id == orders[j].id: flag(orders[i])

O(n²) — a loop inside a loop over the same list. It runs about n²/2 times, and dropping the constant leaves n². This is the exact code from the opening story. The fix is a set: add each id to a set and check membership, turning it into O(n) with O(n) extra memory. Nine minutes becomes a fraction of a second.

2. Two loops, one after the other
for user in users: send_welcome(user) for user in users: log_signup(user)

O(n) — not O(n²). Sequential loops add: n + n = 2n, and the constant 2 gets dropped. Merging them into one loop makes the code twice as fast in wall-clock terms but leaves the complexity class unchanged. That distinction — real speedup, same Big O — is worth internalising.

3. Halving the search space
lo, hi = 0, len(sorted_items) - 1 while lo <= hi: mid = (lo + hi) // 2 if sorted_items[mid] == target: return mid if sorted_items[mid] < target: lo = mid + 1 else: hi = mid - 1

O(log n) — binary search. Each pass discards half of what’s left, so a billion items takes about 30 passes. The catch is the precondition: the list must already be sorted, and sorting costs O(n log n). Sorting once to search many times is a great trade; sorting once to search once is not.

4. A nested loop that isn’t quadratic
for order in orders: # n orders for tax in TAX_BANDS: # always 7 bands apply(order, tax)

O(n) — linear, despite the nesting. The inner loop’s length is fixed at 7 and never grows with the input, so the total is 7n, and the 7 is a constant. Nesting only multiplies complexity when both loops scale with the input. Judge by what grows, not by indentation.

5. The classic recursive trap
def fib(n): if n <= 1: return n return fib(n - 1) + fib(n - 2)

O(2ⁿ) — each call spawns two more, so the call tree roughly doubles at every level. fib(30) makes about 1.6 million calls; fib(50) would run for days. Caching results (memoisation) collapses it to O(n), because each value is then computed exactly once. Same algorithm on paper, a difference of billions in practice.

Best, average and worst case

One algorithm can have three different complexities depending on what you feed it. Searching an unsorted list for a name: if it’s first, you’re done in one step; if it’s last or missing, you check all n.

Best case

The luckiest possible input. Mostly useless for planning — nobody guarantees you luck.

Average case

What typical data does. The most honest number for everyday performance work.

Worst case

The input designed to hurt you. This is what Big O usually reports, and what you should plan around.

The default assumption when someone says “this is O(n log n)” is worst case, unless they say otherwise. That’s deliberately pessimistic, and it’s the right default: your worst case will eventually arrive, and it tends to arrive at the busiest moment. Quicksort is the cautionary tale — O(n log n) on average, O(n²) when the pivots go badly, and the input that triggers it is the very ordinary case of already-sorted data.

Space complexity counts too

Big O describes memory as readily as time. An in-place sort uses O(1) extra space; merge sort needs a scratch array, so O(n). The duplicate-checking fix above trades memory for speed — building a set of every id costs O(n) space to save you from O(n²) time. That trade is almost always worth taking, and recognising when it’s available is most of practical optimisation.

Where this actually matters

Code review

Spotting a nested loop over two growing collections takes seconds and prevents the outage three months later.

Choosing data structures

Array lookup by index is O(1); searching an array is O(n); a hash map turns that search into O(1). Most real speedups are structure changes, not clever code.

Database work

An index converts an O(n) table scan into an O(log n) lookup. That’s the whole reason indexes exist.

Interviews

“What’s the complexity?” is asked in nearly every technical interview, and the follow-up is always “can you do better?”

Don’t over-apply it. If your list has 50 items and always will, an O(n²) loop is completely fine and probably clearer to read. Big O matters when n can grow — and the real skill is asking “how big can this get?” before optimising anything. Most performance work is wasted on code that was never going to be the bottleneck.

Check yourself

Five questions. Open each to check — the correct option is marked.

1. An algorithm takes 4n + 200 steps. What is its Big O?
  • O(4n + 200)
  • O(n)
  • O(200)
  • O(n²)

Drop constants and lower-order terms — only the growth rate survives, and this one grows linearly.

2. Two separate loops over the same list of n items. Total complexity?
  • O(n²)
  • O(n)
  • O(2n²)
  • O(log n)

Sequential loops add — n + n = 2n, which is O(n). Only nested loops multiply.

3. Which class describes binary search on a sorted array?
  • O(1)
  • O(log n)
  • O(n)
  • O(n log n)

Each comparison eliminates half the remaining candidates, so a billion items needs only about 30 steps.

4. Your O(n²) function handles 1,000 records in one second. Roughly how long for 10,000?
  • 10 seconds
  • About 100 seconds
  • 1 second
  • 3 seconds

Ten times the input means a hundred times the work, because the growth is quadratic. This is the calculation to do before shipping, not after.

5. You replace a nested-loop duplicate check with a hash set. What changed?
  • Time O(n²) → O(n), memory unchanged
  • Time O(n²) → O(n), memory O(1) → O(n)
  • Nothing changes
  • Time gets worse, memory improves

You bought a much better time complexity by spending memory on the set. Recognising that trade is most of practical optimisation.

Frequently asked questions

What is Big O notation in simple terms?

It’s a shorthand for how an algorithm’s work grows as its input grows. O(n) means the work grows in step with the data; O(n²) means doubling the data quadruples the work. It describes the shape of that growth, not a measurement in seconds.

Why do we ignore constants in Big O?

Because constants depend on hardware, language and compiler, and they change by small factors — while growth rates change by factors of millions. At a million items the difference between 3n and n is trivial; the difference between n and n² is a trillion steps.

Is O(1) always faster than O(n)?

Not necessarily at small sizes. A constant-time operation with heavy overhead can lose to a linear scan of ten items. Big O describes behaviour as n grows large, so it’s a statement about trends, not a guarantee about any particular input.

What’s a good Big O to aim for?

O(1) and O(log n) are excellent, O(n) is usually fine and often unavoidable, O(n log n) is the practical target for sorting. Treat O(n²) as a warning sign worth investigating, and anything exponential as unusable beyond tiny inputs.

Do I need Big O if I’m not doing interviews?

Yes, though you’ll use it informally. You don’t need the formal maths, but you do need the instinct that says “this loop is inside that loop, and both grow with the data” — that instinct is what stops a working feature becoming an outage once real data arrives.

The takeaway

Big O answers one question: as your data grows, does this code degrade gently or catastrophically? O(1) and O(log n) barely notice. O(n) and O(n log n) scale honestly. O(n²) and worse are fine on toy data and fatal on real data.

You don’t need to derive it formally. You need to look at a function and ask two questions: what grows here, and is anything nested inside anything else that also grows? That habit catches the overwhelming majority of real performance problems, long before they reach production.

Try it on something you wrote this week. Find your longest loop, ask what its input can realistically grow to, and check whether anything scaling sits inside it. That’s the whole practice — and it takes about thirty seconds once the six shapes above are familiar.

big o notationtime complexityalgorithmsspace complexityperformancecomputer science

The post Big O Notation Explained with Real Examples appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/big-o-notation-explained/feed/ 0 842
Sorting Algorithms Compared: Bubble vs Merge vs Quick https://learnwithexamples.org/bubble-vs-merge-vs-quick/ https://learnwithexamples.org/bubble-vs-merge-vs-quick/#respond Tue, 04 Aug 2026 16:14:19 +0000 https://learnwithexamples.org/?p=834 Learn With Examples · Computer Science Three ways to put a list in order. One you’d never ship, one you can trust with anything, and one that beats them both…

The post Sorting Algorithms Compared: Bubble vs Merge vs Quick appeared first on Learn With Examples.

]]>

Learn With Examples · Computer Science

Three ways to put a list in order. One you’d never ship, one you can trust with anything, and one that beats them both — until the day it doesn’t. Here’s how each actually works, watched step by step.

Hand a deck of shuffled cards to three people and watch what they do. The first goes along the row swapping any two neighbours in the wrong order, over and over, until a full pass changes nothing. The second splits the deck in half, sorts each half, then merges the two sorted piles by repeatedly taking the smaller of the two top cards. The third picks one card, throws everything smaller to the left and everything bigger to the right, then repeats on each side.

Those three people just performed bubble sort, merge sort and quicksort. No pseudocode required — the ideas are that physical. What separates them is not cleverness but how many comparisons they need, and that difference explodes as the deck gets bigger.

This article walks through all three step by step, then answers the question that actually matters at work: which one do you reach for, and when does the popular choice betray you?

The 20-second version

Bubble sort compares neighbours and swaps them. Beautifully simple, quadratic time, useless beyond a few hundred items. Learn it, then never ship it.

Merge sort splits, sorts, and merges. Reliably O(n log n) in every case and stable, but needs extra memory the size of your list.

Quicksort partitions around a pivot. Usually the fastest of the three in practice and sorts in place, but a bad pivot can drag it down to quadratic.

Step through all three

Pick an algorithm, then tap through the numbered steps. Cyan marks a comparison, pink a value being moved, amber quicksort’s pivot. All three start from the same six numbers — count how many steps each one needs.

Step through all three

tap the numbers

Step 1 of 15 — 5 was bigger than 1 — swap them

Step 2 of 15 — 5 was bigger than 4 — swap them

Step 3 of 15 — 5 was bigger than 2 — swap them

Step 4 of 15 — 5 and 8 are already in order — leave them

Step 5 of 15 — 8 was bigger than 3 — swap them

Step 6 of 15 — 1 and 4 are already in order — leave them

Step 7 of 15 — 4 was bigger than 2 — swap them

Step 8 of 15 — 4 and 5 are already in order — leave them

Step 9 of 15 — 5 was bigger than 3 — swap them

Step 10 of 15 — 1 and 2 are already in order — leave them

Step 11 of 15 — 2 and 4 are already in order — leave them

Step 12 of 15 — 4 was bigger than 3 — swap them

Step 13 of 15 — 1 and 2 are already in order — leave them

Step 14 of 15 — 2 and 3 are already in order — leave them

Step 15 of 15 — A full pass with no swaps — the list is sorted

Step 1 of 22 — Merging the sorted pieces [1] and [4]

Step 2 of 22 — Smaller front card is 1 — write it into slot 2

Step 3 of 22 — One pile is empty — copy 4 across

Step 4 of 22 — Merging the sorted pieces [5] and [1, 4]

Step 5 of 22 — Smaller front card is 1 — write it into slot 1

Step 6 of 22 — Smaller front card is 4 — write it into slot 2

Step 7 of 22 — One pile is empty — copy 5 across

Step 8 of 22 — Merging the sorted pieces [8] and [3]

Step 9 of 22 — Smaller front card is 3 — write it into slot 5

Step 10 of 22 — One pile is empty — copy 8 across

Step 11 of 22 — Merging the sorted pieces [2] and [3, 8]

Step 12 of 22 — Smaller front card is 2 — write it into slot 4

Step 13 of 22 — One pile is empty — copy 3 across

Step 14 of 22 — One pile is empty — copy 8 across

Step 15 of 22 — Merging the sorted pieces [1, 4, 5] and [2, 3, 8]

Step 16 of 22 — Smaller front card is 1 — write it into slot 1

Step 17 of 22 — Smaller front card is 2 — write it into slot 2

Step 18 of 22 — Smaller front card is 3 — write it into slot 3

Step 19 of 22 — Smaller front card is 4 — write it into slot 4

Step 20 of 22 — Smaller front card is 5 — write it into slot 5

Step 21 of 22 — One pile is empty — copy 8 across

Step 22 of 22 — Every level merged — the list is sorted

Step 1 of 17 — Pivot is 3 — everything smaller goes left of it

Step 2 of 17 — 5 is not smaller than 3 — leave it

Step 3 of 17 — 1 is smaller than 3 — move it left

Step 4 of 17 — 4 is not smaller than 3 — leave it

Step 5 of 17 — 2 is smaller than 3 — move it left

Step 6 of 17 — 8 is not smaller than 3 — leave it

Step 7 of 17 — Pivot 3 drops into its final position, permanently

Step 8 of 17 — Pivot is 2 — everything smaller goes left of it

Step 9 of 17 — Pivot 2 drops into its final position, permanently

Step 10 of 17 — Pivot is 4 — everything smaller goes left of it

Step 11 of 17 — 5 is not smaller than 4 — leave it

Step 12 of 17 — 8 is not smaller than 4 — leave it

Step 13 of 17 — Pivot 4 drops into its final position, permanently

Step 14 of 17 — Pivot is 5 — everything smaller goes left of it

Step 15 of 17 — 8 is not smaller than 5 — leave it

Step 16 of 17 — Pivot 5 drops into its final position, permanently

Step 17 of 17 — Every pivot placed — the list is sorted

Same starting list every time: [5, 1, 4, 2, 8, 3]. Count the steps each algorithm needs — that gap is the whole argument.

Bubble sort: the one everyone learns first

Bubble sort does exactly one thing: walk the list, compare each pair of neighbours, swap them if they’re out of order. Repeat until a full pass produces no swaps at all. Large values “bubble” to the end one position per pass, which is where the name comes from.

Sort [5, 1, 4, 2] by hand:

PassComparisonList after
15 vs 1 → swap[1, 5, 4, 2]
15 vs 4 → swap[1, 4, 5, 2]
15 vs 2 → swap[1, 4, 2, 5]
21 vs 4 → keep[1, 4, 2, 5]
24 vs 2 → swap[1, 2, 4, 5]
3no swaps made[1, 2, 4, 5] ✓
// bubble sort with the early-exit optimisation function bubbleSort(a) { for (let i = 0; i < a.length - 1; i++) { let swapped = false; for (let j = 0; j < a.length - 1 - i; j++) { if (a[j] > a[j + 1]) { [a[j], a[j + 1]] = [a[j + 1], a[j]]; swapped = true; } } if (!swapped) break; // already sorted, stop early } return a; }

The cost is brutal. Every pass compares nearly every pair, and you need close to n passes, so the work grows with . Ten items cost about 45 comparisons. A thousand items cost roughly half a million. Ten thousand cost fifty million. You can feel that in a browser tab.

The one situation where bubble sort isn’t embarrassing is nearly sorted data. With the early-exit check above, an already-sorted list is verified in a single pass of n-1 comparisons — that’s O(n), better than merge sort’s best case. If you have a list where one element occasionally drifts out of place, a bubble pass is a perfectly reasonable repair. That’s a narrow niche, and insertion sort usually fills it better, but it’s real.

Why teach it at all, then? Because it makes the shape of the problem visible. Once you’ve watched bubble sort waste 20 comparisons re-checking pairs it already knows are fine, the motivation behind divide-and-conquer stops being abstract. It’s the algorithm you learn in order to want a better one.

Merge sort: split, sort, stitch

Merge sort refuses to compare distant items at all. Instead it breaks the list in half, again and again, until every piece has a single element — and a single element is sorted by definition. Then it walks back up, merging pairs of sorted pieces.

The merge step is where the magic sits, and it’s the part worth understanding physically. You have two sorted piles face up. Look at the top card of each. Take the smaller one. Repeat. Because both piles are sorted, the smallest remaining card is always one of the two you’re looking at — you never search, you never backtrack.

Merging [1, 4] and [2, 3]:

Left pileRight pileCompareOutput
[1, 4][2, 3]1 vs 2 → take 1[1]
[4][2, 3]4 vs 2 → take 2[1, 2]
[4][3]4 vs 3 → take 3[1, 2, 3]
[4][ ]right empty → copy 4[1, 2, 3, 4]

Count the levels: halving a list of 1,000 takes about 10 splits to reach single items, because 2¹⁰ is 1,024. Each level does roughly n comparison work to merge everything back. Ten levels × a thousand items ≈ 10,000 operations, against bubble sort’s half a million. That’s the entire argument for O(n log n), and it’s why merge sort’s worst case, best case and average case are all the same — the splitting doesn’t care what the data looks like.

Two properties make merge sort the professional’s safe choice:

It’s stable

Equal items keep their original relative order. Sort orders by date, then by customer, and the date order survives inside each customer. Quicksort does not promise this.

It’s predictable

No input can make it slow. For latency-sensitive systems, a guaranteed ceiling beats a lower average with a nasty tail.

It works off-disk

Merging needs only the front of each pile, so you can sort a 500 GB file with 8 GB of RAM. This is how external sorting works.

It costs memory

The standard version needs a scratch array as big as the input. On memory-tight systems, that’s the dealbreaker.

Quicksort: pick a pivot, throw everything to one side

Quicksort also divides and conquers, but it partitions before recursing rather than after. Choose one element as the pivot. Rearrange the list so everything smaller sits left of it and everything larger sits right. The pivot is now in its final position, permanently. Then repeat on the left chunk and the right chunk.

Take [7, 2, 9, 4, 5] with 5 as the pivot. Walk the rest: 7 is bigger (right), 2 is smaller (left), 9 is bigger (right), 4 is smaller (left). You get [2, 4] 5 [7, 9]. One pass, and 5 is done forever. Now solve the two small pieces the same way.

// quicksort, Lomuto partition scheme function quickSort(a, lo = 0, hi = a.length - 1) { if (lo >= hi) return a; const pivot = a[hi]; let i = lo; for (let j = lo; j < hi; j++) { if (a[j] < pivot) { [a[i], a[j]] = [a[j], a[i]]; i++; } } [a[i], a[hi]] = [a[hi], a[i]]; // pivot into place quickSort(a, lo, i - 1); quickSort(a, i + 1, hi); return a; }

On average the pivot lands somewhere near the middle, the problem halves each time, and you get O(n log n) — typically with a smaller constant factor than merge sort, because quicksort swaps within the original array instead of copying into a scratch one. Fewer memory writes, better cache behaviour, faster in the real world.

The trap. Suppose you always pick the last element as pivot, and the list is already sorted. Every pivot is the largest remaining value, so one side gets everything and the other gets nothing. You’ve turned O(n log n) into O(n²) — and the input that triggers it is the most common input in the world: data that’s already in order. Real implementations dodge this by picking a random pivot, or the median of the first, middle and last elements.

Race them on the same data

Same shuffled array, three algorithms, one operation per tick. This isn’t a wall-clock benchmark — it’s an operations count, which is what the big-O notation is actually measuring.

Head-to-head race

tick to start
Bubble ~1,800 comparisons
Merge ~300 comparisons
Quick ~250 comparisons

Each bar advances at a speed proportional to the work its algorithm really does on a shuffled 60-item list. Quick and merge finish while bubble sort is still grinding through its second pass.

On random data quicksort usually finishes first, merge close behind, bubble far back. Hand all three an already-sorted list, though, and the ranking inverts completely: bubble sort exits after a single clean pass, while quicksort with a naive last-element pivot collapses into its worst case. Same three algorithms, opposite result, purely because the input changed shape.

Big-O tells you how an algorithm behaves as data grows. It doesn’t tell you which one wins on your data. Only the shape of your input decides that.

The numbers, side by side

BubbleMergeQuick
Best caseO(n)O(n log n)O(n log n)
AverageO(n²)O(n log n)O(n log n)
Worst caseO(n²)O(n log n)O(n²)
Extra memoryO(1)O(n)O(log n)
Stable?YesYesNo
In place?YesNoYes
Use it whenTeaching, or tiny nearly-sorted listsYou need guarantees, stability, or external sortingYou want raw speed on in-memory data

Those symbols get abstract fast, so put real numbers on them. Pick a list size and watch the gap between quadratic and logarithmic growth open up.

How bad does n² get?

pick a list size
45Bubble ops (n²)
33Merge / quick ops
1.4×Times more work
instantBubble time est.
5,000Bubble ops (n²)
664Merge / quick ops
7.5×Times more work
under 1 msBubble time est.
500,000Bubble ops (n²)
9,966Merge / quick ops
50×Times more work
5 msBubble time est.
5 billionBubble ops (n²)
1,660,964Merge / quick ops
3,010×Times more work
50 secondsBubble time est.

Notice how the ratio behaves. At 100 items bubble sort does around 15 times more work — annoying, survivable. At 100,000 items it does over 3,000 times more. That is the difference between a page that renders instantly and one that hangs the browser for a minute. Complexity classes don’t matter much at small scale and matter enormously at large scale, which is exactly why beginners under-rate them.

So which one do you actually use?

Honest answer for day-to-day work: call your language’s built-in sort. It’s been tuned by people who do nothing else. But knowing what’s under it tells you when to override it.

Sorting objects by two fields

You need stability — merge sort, or a built-in that guarantees it. Sort by the secondary key first, then the primary.

Big array, memory is tight

Quicksort. In-place, cache-friendly, only recursion stack overhead. Just randomise the pivot.

Data bigger than RAM

Merge sort. It’s the only one of the three that sorts chunks on disk and merges streams.

Worst case must be bounded

Merge sort. Real-time and latency-critical systems care about the ceiling, not the average.

Fewer than ~20 items

Use insertion sort. Its overhead is so low it beats the clever algorithms at small sizes — which is why real implementations switch to it below a threshold.

Explaining sorting to someone

Bubble sort, once. Then show them the step-through above and let the operation count make the argument.

What most standard libraries actually run is a hybrid. Introsort starts as quicksort, counts its recursion depth, and switches to heapsort if the pivots are going badly — giving quicksort’s speed with a guaranteed O(n log n) ceiling. Timsort, used in Python and Java for objects, is merge sort that detects runs of already-ordered data and exploits them, which makes it startlingly fast on the semi-sorted data that real applications produce. Both are engineering answers to the trade-offs on this page.

Check yourself

Five questions. Open each one to check yourself — the correct option is marked.

1. Why is merge sort’s worst case the same as its best case?
  • It checks the data first
  • The splitting is fixed and doesn’t depend on the values
  • It uses extra memory
  • It’s stable

Merge sort always halves the list, whatever’s in it. The number of levels is fixed at log₂n, so no input can make it slow.

2. Which input makes naive quicksort hit its worst case?
  • Random data
  • Data with many duplicates only
  • Already-sorted data with a last-element pivot
  • Very short lists

Every pivot becomes the largest remaining value, so one partition gets everything. The fix is a random or median-of-three pivot.

3. What does “stable” mean for a sorting algorithm?
  • Equal items keep their original relative order
  • It never crashes
  • It always takes the same time
  • It uses no extra memory

Stability lets you sort by one field, then another, and keep the first ordering inside groups. Quicksort doesn’t guarantee it.

4. Roughly how many comparisons does bubble sort need for 1,000 items?
  • About 1,000
  • About 10,000
  • About 500,000
  • About 1,000,000

Roughly n²/2 — half a million. Merge sort handles the same list in about 10,000.

5. You must sort a 400 GB log file on a machine with 16 GB of RAM. Which approach?
  • Quicksort, it’s in place
  • Merge sort, sorting chunks then merging streams
  • Bubble sort, it uses no extra memory
  • None can do it

External merge sort only needs the front of each sorted run in memory at a time — the classic solution to sorting more data than you can hold.

Frequently asked questions

Which sorting algorithm is fastest?

For general in-memory data, quicksort is usually fastest in practice because it sorts in place with excellent cache behaviour. Merge sort is faster in the worst case, since quicksort can degrade to O(n²) with bad pivots. Bubble sort is slowest by a wide margin at any meaningful size.

Is bubble sort ever useful in real code?

Rarely, but not never. On a list that’s already nearly sorted, the early-exit version runs in O(n) and is trivial to write and verify. For anything else, insertion sort does the same job better and your language’s built-in sort beats both.

Why does quicksort beat merge sort if merge sort has a better worst case?

Constant factors. Quicksort swaps elements inside the original array; merge sort copies into a scratch array and back on every level. Fewer memory writes and better cache locality mean quicksort typically wins on wall-clock time even when both are O(n log n).

What does O(n log n) actually mean in plain terms?

The work grows a little faster than the list size, but nowhere near as fast as squaring it. Double the items and you do slightly more than double the work — whereas O(n²) means doubling the items quadruples the work. That gap is what makes large-scale sorting feasible at all.

Which sort do Python, Java and JavaScript use?

Python uses Timsort, a merge sort variant that detects existing sorted runs. Java uses Timsort for objects and a dual-pivot quicksort for primitives. Most JavaScript engines use Timsort-style stable sorts for Array.sort(). All three are hybrids built from the ideas on this page.

The takeaway

Bubble sort compares neighbours and pays for it quadratically. Merge sort splits the problem into halves whose cost adds up to n log n, every single time, at the price of extra memory. Quicksort partitions around a pivot and is usually the fastest of the three — as long as the pivot is chosen sensibly.

The deeper lesson isn’t which one wins. It’s that the same task can be organised in ways whose costs diverge by a factor of thousands, and that the winning approach depends on the shape of your data, not on the elegance of the code. Step back through the three walkthroughs above and count what each one needed on the very same six numbers — then imagine those gaps at a million items. That is the thing worth remembering.

bubble sortmerge sortquicksortbig-Oalgorithmsdata structures

The post Sorting Algorithms Compared: Bubble vs Merge vs Quick appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/bubble-vs-merge-vs-quick/feed/ 0 834
What Are Tokens and Context Windows? https://learnwithexamples.org/what-are-tokens-and-context-windows/ https://learnwithexamples.org/what-are-tokens-and-context-windows/#respond Tue, 04 Aug 2026 09:21:42 +0000 https://learnwithexamples.org/?p=814 LearnWithExamples A language model never actually sees your sentence. It sees a stream of numbered fragments, and it can only hold so many of them at once. Understand those two…

The post What Are Tokens and Context Windows? appeared first on Learn With Examples.

]]>

LearnWithExamples

A language model never actually sees your sentence. It sees a stream of numbered fragments, and it can only hold so many of them at once. Understand those two facts and almost every strange AI behaviour you have ever run into suddenly has an explanation.

I have been building things on top of language models since the days when “context window” meant 512 fragments and you budgeted them like a family budgets grocery money. The hardware got better. The models got enormous. And yet, in workshop after workshop, the same two questions come up first — and they are almost always asked in the wrong order.

People ask “why did the AI forget what I told it?” before they ask “what exactly is the AI storing in the first place?” That is like asking why your suitcase won’t close before checking what you packed. So let’s do it properly. First we’ll open up a token — see it, split it, count it. Then we’ll build the window it lives inside, and watch what happens when the window fills up.

By the end you’ll be able to look at a prompt and estimate its cost in your head, explain why the model insists “strawberry” has two Rs, know why Hindi or Tamil text costs nearly triple what English does, and — most usefully — know exactly what to trim when a system tells you your input is too long.

The 30-second answer

Tokens and context windows, defined

A token is the smallest chunk of text a language model actually reads. Usually it’s a piece of a word — roughly ¾ of an English word on average. The model converts every token into a number before it does anything else. Text in, numbers out.

A context window is the maximum number of tokens the model can have in front of it at one time — your instructions, the uploaded file, the entire chat history, plus the reply it is currently writing. It’s a hard ceiling, not a soft suggestion. Nothing outside it exists as far as the model is concerned.

Tokens are the units. The context window is the container. Everything else in this article is a consequence of those two sentences.

Part 1 — A token is not a word

Here is the first mental model most beginners build, and it is wrong: “the AI reads word by word, like I do.” Reasonable guess. Also the source of about half the confusion I see.

Think about why splitting on words fails. English alone has well over a million distinct word forms once you include names, typos, plurals, brand names, hashtags, and the thing your colleague types when they mean “definitely.” A model would need a lookup table with a million-plus entries, and the moment somebody typed Kanniyakumari or ngrok or bruhhhh, the table would shrug and return “unknown.”

Now try the opposite extreme: split on characters. Only about a hundred symbols to memorise — clean, complete, nothing is ever unknown. But now the word “encyclopedia” costs twelve slots instead of one, and the model has to learn spelling from scratch before it learns meaning. Sequences get impossibly long, and long sequences are expensive (we’ll get to why).

So the field landed in the middle. Subword tokenization: keep whole tokens for the text fragments that show up constantly, and break rare things into smaller pieces that can be reassembled. Common stuff is cheap. Weird stuff still works, just costs more.

Split by words

Tiny sequences, impossible vocabulary. Any unseen word breaks it. Rejected.

Split by characters

Tiny vocabulary, endless sequences. Nothing breaks, but everything is slow. Rejected.

Split by subwords

Frequent chunks stay whole, rare chunks get split. This is what every major model uses today.

Watch it happen

Reading about tokenization is far less convincing than watching a sentence get shredded. Type into the box below — or hit one of the preset buttons — and see the fragments appear. Each coloured block is one token. Notice the little spaces sitting inside the blocks: in most tokenizers a leading space is part of the token, which is why the and the are two different entries entirely.

Live token splitter

estimate · type to update
0Tokens
0Words
0Characters
0Tokens / word

Grey <byte> blocks are the extra tokens spent on characters outside the basic Latin set. This tool approximates real byte-pair encoding closely enough to build intuition — exact counts vary a little between model families.

How the splitting rules get invented

Nobody sits down and hand-writes the list of fragments. It’s learned from data, most often with an algorithm called Byte Pair Encoding (BPE). The idea is almost embarrassingly simple, and you can run it on paper.

Suppose your entire training corpus is these three words, repeated a lot:

RoundWhat the algorithm seesMost frequent adjacent pairNew merged token
Startl o w · l o w e r · l o w e s tl + olo
2lo w · lo w e r · lo w e s tlo + wlow
3low · low e r · low e s te + rer
4low · low er · low e s te + ses
5low · low er · low es tes + test

Five rounds and the algorithm has independently discovered the root low and the suffixes er and est — real English morphology, learned purely by counting. Run this for millions of rounds on a slice of the internet and you get a vocabulary of roughly 50,000 to 200,000 fragments, ordered by how useful they are.

Two consequences fall straight out of this, and both matter in practice:

  • Frequency decides price. Text that looked like the training data gets packed efficiently. Text that didn’t gets shredded into crumbs. Your name is probably one token if you’re called David and five tokens if you’re called Devanshi.
  • The vocabulary is frozen at training time. A model trained mostly on English carries an English-shaped tokenizer forever. It can still read Hindi, Swahili, or Korean — it just pays a heavier toll per sentence.
Example you can feel

Type “Learning is fun” into the splitter above, then type the same sentence in Hindi. English: about 3–4 tokens. Hindi: often 15–25. Same meaning, four to six times the cost. If you build a product for an Indian audience and price it per token, this single detail can decide whether your unit economics work.

Part 2 — Token maths you can do in your head

You will rarely need an exact count. You will constantly need a fast estimate: will this fit? These are the ratios I have used for years, and they hold up well enough for planning.

Rule of thumbApproximate valueHow to use it
English words → tokens1 word ≈ 1.3 tokensMultiply your word count by 4/3.
Characters → tokens4 characters ≈ 1 tokenFastest check: divide character count by 4.
A typical page of prose≈ 600 tokens450 words per double-spaced page.
Source code1 line ≈ 8–12 tokensIndentation and punctuation are surprisingly costly.
Meeting transcript1 minute ≈ 200 tokensSpeech runs near 150 words per minute.
Non-Latin scripts2–4× the English costDevanagari, Tamil, Thai, Arabic, Korean.
Emoji1 emoji ≈ 2–4 tokensSkin-tone and family emoji are the worst offenders.

Let’s apply them to things you actually handle:

A WhatsApp message

~25 words → ≈ 33 tokens. You could send 6,000 of these before filling a 200K window.

A one-page résumé

~600 words → ≈ 800 tokens. Screening 100 résumés at once is genuinely feasible.

A 12-page contract

~6,000 words → ≈ 8,000 tokens. Comfortable in any modern window.

A 90-minute meeting

~13,500 words → ≈ 18,000 tokens. Fine alone; painful if you also attach the slides.

A 300-page novel

~100,000 words → ≈ 133,000 tokens. Fits in a large window, with room to spare for questions.

A mid-size codebase

40,000 lines → ≈ 400,000 tokens. This is where retrieval stops being optional.

Part 3 — The five places tokenization gets weird

This is the part of the article I’d tell my past self to read first. Almost every “the AI is stupid” complaint I have investigated turned out to be a tokenizer artefact, not a reasoning failure.

1. The model cannot reliably count letters

Ask a model how many times the letter R appears in “strawberry” and you may get “two.” People post screenshots of this as proof that AI is a parlour trick. What actually happened is subtler: the model never received the letters. It received something like str + aw + berry — three opaque ID numbers. Asking it to count Rs is like asking you to count the letters in a word someone spelled out in Morse code, at speed, without writing anything down. Not impossible, but a genuinely awkward task given the input.

Fix it in practice

Force the letters apart before asking: s-t-r-a-w-b-e-r-r-y. Hyphenating breaks the word into single-character tokens and accuracy jumps immediately. Same trick works for reversing words, counting syllables, and checking palindromes.

2. Numbers split in unhelpful places

Depending on the tokenizer, 148392 might become 148+392, or 1+48+39+2. The digits do not line up in neat place-value columns the way they do on paper, which is a large part of why models were historically shaky at long arithmetic. It also means an invoice number, a PAN, an order ID, or a phone number can be far more expensive than its length suggests — and is easier to mis-copy.

3. A trailing space quietly poisons your prompt

If your prompt ends with "The answer is " — with a space after “is” — you have handed the model a dangling space token and asked it to continue. But the natural next token is " 42", which already includes its own leading space. You’ve created a conflict the model must work around, and quality drops for no visible reason. Fifteen years in, and I still occasionally lose twenty minutes to a stray space at the end of a template string.

4. Whitespace is a real line item in code

Deeply indented code, four-space tabs, blank lines between functions — every bit of that costs tokens. A well-formatted 500-line Python file can cost noticeably more than the same logic minified. This is not an argument for writing ugly code; it is an argument for sending the model the three functions that matter instead of the whole file.

5. Rare and mixed-script text costs a fortune

Scientific nomenclature, chemical formulas, Sanskrit terms, transliterated place names, base64 blobs, UUIDs, and minified JavaScript all tokenize horribly. A single UUID can eat 20+ tokens. If you are pasting logs full of them, you are burning most of your budget on strings the model does not need to read carefully.

Every unusual character you send is a small tax. Every ordinary English word is a discount. That asymmetry is baked into the tokenizer, and no amount of prompt engineering removes it.

Part 4 — The context window: the desk, not the brain

Now the second half. You know what tokens are; the question is how many the model can look at simultaneously.

My favourite analogy comes from a colleague who used to work in an architecture office. Picture a draughtsman at a drawing board. The board is a fixed size. Everything he needs — the client brief, the site survey, last week’s revisions, the sheet he is drawing on right now — has to be laid out on that board at once. Anything that doesn’t fit goes on the floor. And here is the crucial part: he cannot remember what’s on the floor. Not “remembers it vaguely.” Cannot see it at all.

That board is the context window. It is measured in tokens, and it holds four things simultaneously:

ONE CONTEXT WINDOW — 100% OF WHAT THE MODEL CAN SEE Systempromptrules, persona Attached files & retrieved docsPDFs, database rows, search results,tool outputs, images-as-tokens Conversation historyevery earlier message from both sides,resent in full on every single turn Reserved forthe replyoutput tokens ANYTHING THAT DOESN’T FIT Dropped, truncated, or summarised away The model has no awareness that this content ever existed. It cannot ask for it back.
The four tenants of every context window. They compete for the same fixed space — grow one and you shrink the others.

Read that diagram twice, because it contains the single most misunderstood fact about chatbots: the conversation history is re-sent, in full, on every turn. The model is not sitting there remembering you between messages. It is stateless. Each time you press enter, the application quietly bundles up the entire transcript so far, appends your new message, and ships the whole parcel off again. That is why turn 40 of a conversation costs far more than turn 2, and why a long chat gradually slows down.

Watch a conversation fall off the edge

The simulator below runs a customer-support chat. Drag the slider to shrink or grow the window and watch which messages survive. The system prompt is pinned — it is always sent — and the newest messages have priority. Everything greyed out is on the floor.

Context window simulator

drag the slider
system prompt history in window free space 0 messages dropped

Shrink the window far enough and the model loses the customer’s order number — then confidently asks for it again, or worse, invents one. That “sudden amnesia” moment is not a bug in the model. It is arithmetic.

Part 5 — Who is eating your window?

When a team tells me “we have a 128,000 token window, we’ll never hit it,” I ask them to write down the budget. It is always tighter than they expect. Here’s a real-shaped example — an internal HR assistant:

Line itemTokensNotes
System prompt (tone, rules, escalation policy)1,400Sent on every single turn, forever.
Tool / function definitions2,100Six tools with parameter schemas.
Few-shot examples1,800Four demonstration Q&A pairs.
Retrieved policy documents (8 chunks)6,400The RAG payload.
Employee’s uploaded payslip PDF3,200Tables tokenize badly.
Conversation history (turn 18)9,700Grows every turn, never shrinks by itself.
Reserved for the reply2,000If you forget this, generation gets cut off mid-sentence.
Total26,600And this is a simple assistant.

Notice what dominates: not the user’s question, which is maybe 30 tokens. The scaffolding dominates. That is the normal state of affairs, and it’s why the highest-leverage optimisation is almost never “write a shorter question.”

The reservation nobody makes

Input and output share the same window. If you fill 199,000 of a 200,000-token window with a document and then ask for a detailed 5-page report, there is no room left to write it in. The response gets truncated mid-sentence and everyone blames the model. Always leave headroom for the answer — I budget 10–15% by default.

Input tokens and output tokens are not the same product

Every provider charges more per output token than per input token, often 3–5×, and beginners assume this is arbitrary. It isn’t. Reading your input happens in one parallel pass — the whole prompt goes through the model at once. Writing the reply happens one token at a time, sequentially, with a full pass through the network for each token produced. Output is genuinely the expensive half.

The practical takeaway: a request that reads 50,000 tokens and returns a 200-token summary is cheap. A request that reads 500 tokens and generates a 4,000-word essay may cost more. If your bill is climbing, look at what you are asking the model to write, not just what you are feeding it.

What actually fits in a window?

pick a size
0English words
0Pages of prose
0Lines of code
0Minutes of transcript
0Business emails
0Hindi words

Part 6 — Why bigger windows didn’t solve everything

Around the time million-token windows arrived, a lot of people declared retrieval dead. Just dump everything in! It did not play out that way, for four reasons I now explain in every architecture review.

Attention cost grows faster than the input

In a standard transformer, every token compares itself with every other token. Double the input and you roughly quadruple that work. Modern implementations soften the curve considerably, but the shape of the problem remains: long contexts are disproportionately expensive in compute and in memory. That cost lands on you as latency and as money.

Information gets lost in the middle

Researchers have repeatedly measured a U-shaped accuracy curve: models are excellent at recalling material near the beginning of the context and near the end, and noticeably weaker on material buried in the middle. Put the critical clause on page 40 of a 90-page contract and recall genuinely suffers. It is the reading equivalent of remembering the first and last speaker at a conference and losing the eleven in between.

Two habits that follow from this

First: put your instructions after the long document, not before it — the tail position is a strong one. Second: for anything critical, state it twice, once at the top and once at the bottom. It feels redundant. It measurably works.

More context means more distraction

Give a model 50 pages when the answer lives in one paragraph, and you have added 49 pages of plausible-looking, semantically similar noise. Precision drops. I have watched a support bot’s accuracy improve by 12 points simply by retrieving four document chunks instead of twenty. Less, but better, beat more.

You pay for every token, every turn

A 300,000-token prompt sent on each of 20 conversational turns is six million input tokens for a single conversation. Multiply by a thousand users. This is the bill that surprises startups in month three.

A large context window is a bigger desk, not a better draughtsman. What you place on the desk still decides the quality of the drawing.

Part 7 — What happens when you overflow

Four different things, and knowing which one your tool does explains a lot of otherwise baffling behaviour.

Hard rejection

The API refuses the request outright with a “maximum context length exceeded” error. Honest and annoying — typical of direct API calls.

Silent truncation

The oldest content is chopped off without telling anyone. The model answers confidently using half the information. This is the dangerous one.

Sliding window

Old turns drop off the back as new ones arrive, like a conveyor belt. Chat apps often do this. It’s why long chats “forget” the beginning.

Rolling summarisation

Older turns get compressed into a paragraph of notes that stays in context. Cheapest in tokens, but detail is permanently lost.

Here’s the failure I see most often in production, told as a small story. A logistics company built an assistant that read a shipment manifest and answered questions about it. It worked beautifully in testing. In production it started confidently reporting wrong container weights. The cause: their largest customers had manifests that pushed past the window, their framework silently truncated the head of the file — including the column headers — and the model, doing its best with headerless numbers, guessed which column was which. No error was ever logged. Nobody had built a token counter into the pipeline.

Count your tokens before you send them. That is the whole lesson. It takes four lines of code and it prevents an entire genus of bug.

Part 8 — The practical playbook

Nine tactics, roughly in the order I reach for them.

TacticWhat you actually doBest when
Chunk and retrieveSplit documents into 300–800 token passages, embed them, and pull only the handful that match the question.Your corpus is bigger than any window — knowledge bases, manuals, archives.
Summarise and carryEvery ~10 turns, compress the conversation into a short state note and drop the raw transcript.Long-running assistants and agents.
Sandwich your instructionsState the task before the document and repeat it after.Any prompt over ~10,000 tokens.
Reserve output spaceCap input at 85% of the window; leave the rest for the reply.Always. Non-negotiable.
Strip the noiseRemove HTML tags, boilerplate footers, base64 blobs, repeated log timestamps, license headers.Scraped web content and log files — often a 40% saving.
Keep the stable parts firstPut the system prompt and fixed examples at the very start, variable content after.You’re using prompt caching — cache hits need an identical prefix.
Ask for structureRequest JSON or a bounded list instead of free prose.Output tokens are your cost driver.
Split the jobThree focused calls beat one giant call that tries to do everything.Multi-step analysis; also improves accuracy.
Right-size the modelRoute simple classification to a small model, reserve the big window for the hard step.Cost is out of control and latency matters.

A worked example: 200 support tickets

Say you want themes across 200 support tickets averaging 400 words each. Naive approach: paste all of them, about 107,000 tokens, into one prompt and ask for themes. It fits in a large window — and produces a bland, shallow answer, because the middle 150 tickets get skimmed.

What actually works: batch the tickets in groups of 20. For each batch, ask for a compact structured summary — five themes, each with a one-line description and a count. That’s ten calls of ~11,000 input tokens each, returning ~300 tokens apiece. Then a final call takes those ten summaries (3,000 tokens total) and merges them into the master list. Total input is similar, output is small, and every ticket got real attention rather than a glance. This map-then-reduce shape is the workhorse pattern of practical LLM engineering, and it exists entirely because of context limits.

Part 9 — Five myths worth deleting

“The model remembers our previous conversations.”

By default, no. Each request is stateless. If a product appears to remember you across sessions, an application layer is storing notes in a database and quietly pasting the relevant ones into the context window before the model ever sees your message. The memory lives in the product, not in the model.

“A bigger context window means the AI is smarter.”

Window size measures capacity, not capability. A model with a million-token window can still reason worse than a model with 32,000. They’re independent specifications — like confusing the size of a desk with the skill of the person sitting at it.

“One token is one word.”

Only by coincidence, and mostly for short common English words. Average English runs about 1.3 tokens per word; a rare surname might be five tokens; a Devanagari sentence can be three or four times its English equivalent. The splitter above makes this obvious in ten seconds.

“Images and files don’t use tokens.”

They absolutely do. Images are converted into token-equivalents based on their dimensions — a high-resolution screenshot can cost well over a thousand. PDFs are extracted to text and tokenized like anything else, and tables and scanned pages are especially expensive. Audio and video are worse still.

“If it fits in the window, the model reads it all equally.”

Fitting is necessary, not sufficient. Attention is uneven — beginnings and endings dominate. Fitting a document into context guarantees the model can see it, never that it will weigh every paragraph the same.

Quick reference

TermIn one line
TokenThe smallest unit of text a model reads; usually a word fragment.
TokenizerThe component that converts text into token IDs and back again.
VocabularyThe fixed list of all tokens a model knows — typically 50K–200K entries.
BPEByte Pair Encoding — the merge-the-most-frequent-pair algorithm that builds that list.
Context windowMaximum tokens the model can hold at once, input and output combined.
Input tokensEverything you send: prompt, history, files, tool definitions.
Output tokensEverything the model writes back. More expensive, generated one at a time.
TruncationCutting content that doesn’t fit — sometimes silently.
RAGRetrieval-augmented generation: fetch only the relevant passages, then answer.
Prompt cachingReusing the processed form of an unchanged prompt prefix to cut cost and latency.
KV cacheStored intermediate values for tokens already processed; grows with context length.
Lost in the middleThe measured tendency to recall the start and end of a long context better than the middle.

Check yourself

Five questions. No score is recorded anywhere — this is just for you.

Frequently asked questions

How many tokens is 1,000 words?

About 1,300 tokens for ordinary English prose. Technical writing with lots of code, names, or numbers runs higher — budget 1,500. Hindi, Tamil, Arabic or Korean text of the same length can reach 3,000–5,000.

How do I count tokens exactly?

Use the tokenizer library published for your specific model family — the count is model-specific, so a number from one vendor’s tool won’t match another’s precisely. For planning, the four-characters-per-token estimate is close enough; for billing and hard limits, count properly in code before you send.

Does a longer chat cost more even if my messages are short?

Yes, and this surprises almost everyone. The full history is re-sent every turn, so input cost grows roughly with the square of conversation length. Starting a fresh chat for a new topic is not just tidier — it is cheaper and usually produces better answers.

What’s the difference between context window and training data?

Training data is what shaped the model’s weights months ago — permanent, general, and unchangeable at chat time. The context window is working memory for right now — temporary, specific, and gone when the session ends. A model can be trained on the whole internet and still not know what’s in a document you haven’t pasted.

Can I make my text use fewer tokens?

Yes, meaningfully. Strip boilerplate and markup, drop repeated headers and footers, replace long IDs with short labels, summarise older conversation turns, send excerpts instead of whole files, and prefer plain common vocabulary over ornate phrasing. Compressing prompts by 30–50% with no loss of meaning is routine once you look for it.

Why do models struggle with counting letters and long arithmetic?

Because tokenization hides the internal structure. The model sees IDs for word fragments, not individual characters or place-value digits. Spacing the characters out (s-t-r-a-w) or asking it to use a calculator tool both fix it, and both work because they change what the tokenizer produces.

Where this leaves you

Two ideas, and everything downstream follows from them. Text becomes tokens — fragments, not words, priced by how ordinary they are. Tokens live inside a window — a hard-edged desk, not a memory, holding your instructions, your documents, your entire chat history and the answer being written, all at the same time.

Once those two things are solid, the rest of the field stops looking like magic. Retrieval is a technique for choosing what goes on the desk. Prompt caching is a technique for not re-reading the same corner of it. Summarisation buys space. Chunking buys attention. Even the odd failures — the miscounted letters, the sudden amnesia at message 40, the report that stops mid-sentence — become predictable, and therefore preventable.

Try this before you close the tab: take a prompt you use regularly, paste it into the splitter near the top of this page, and look at the number. Most people discover their standard prompt is two or three times more expensive than they assumed, and that a third of it is boilerplate the model never needed. That’s a five-minute experiment that has saved teams I’ve worked with real money — and it starts with just looking at the tokens.

tokenscontext windowtokenizationBPEprompt engineeringLLM basics

The post What Are Tokens and Context Windows? appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/what-are-tokens-and-context-windows/feed/ 0 814
The Fundamental Theorem of Calculus — Explained with Real Examples https://learnwithexamples.org/fundamental-theorem-of-calculus/ https://learnwithexamples.org/fundamental-theorem-of-calculus/#respond Mon, 27 Jul 2026 16:43:01 +0000 https://learnwithexamples.org/?p=804 The Fundamental Theorem of Calculus — Explained with Real Examples Learn · With · Examples A car’s speedometer and odometer are secretly the two halves of the most important theorem…

The post The Fundamental Theorem of Calculus — Explained with Real Examples appeared first on Learn With Examples.

]]>
The Fundamental Theorem of Calculus — Explained with Real Examples

Learn · With · Examples

A car’s speedometer and odometer are secretly the two halves of the most important theorem in calculus. One measures a rate, the other accumulates it — and the Fundamental Theorem of Calculus is simply the precise statement of how differentiation and integration undo each other. This guide builds it from that one everyday example, with interactive demos, worked problems, and code.

📖 20 min read 📈 FTC Part 1 explorer 📐 FTC Part 2 explorer 🎬 Area accumulation demo ❓ 6-question quiz

Section 01

The Big Idea — Two Operations, One Relationship

For most of calculus, derivatives and integrals feel like two completely separate tools. A derivative measures how fast something is changing at an instant — a slope. An integral measures the total accumulation of something over an interval — an area. They seem to answer opposite kinds of questions.

The Fundamental Theorem of Calculus (FTC) is the discovery that they are not separate at all. They are inverse operations — the same relationship that addition has with subtraction, or multiplication has with division. Differentiate an integral and you get back the original function. Integrate a derivative and you get back the original function (up to a constant). This single fact is why calculus works as one unified subject instead of two disconnected topics.

💡 The Key Idea

Every accumulation (integral) has an accumulation rate (derivative), and that rate is just the original function you were accumulating. Differentiation undoes integration. That’s the entire theorem — everything else is precise bookkeeping around that one idea.

2
Parts to the theorem — one about derivatives of integrals, one about evaluating integrals
1670s
Decade Newton and Leibniz independently formalized this connection
F(b)−F(a)
The entire computational shortcut FTC Part 2 hands you
Rectangles you’d need to sum by hand without this theorem — FTC replaces that with 2 substitutions

Section 02

Formal Statement — FTC Part 1 and Part 2

If F(x) = ∫[a to x] f(t) dt,   then F'(x) = f(x)
FTC PART 1 — the derivative of an accumulation function gives back the original function
∫[a to b] f(x) dx = F(b) − F(a)
FTC PART 2 — where F is ANY antiderivative of f (F’ = f)

Part 1 tells you that integration and differentiation are inverse processes. Part 2 is the practical payoff: it hands you a way to compute the exact value of a definite integral without summing infinitely many rectangles — just find an antiderivative and plug in the two endpoints.

⚠ Continuity Requirement

Both parts require f to be continuous on the relevant interval. If f has a jump, a hole, or shoots to infinity somewhere in [a,b], the theorem doesn’t directly apply without extra care — this is why calculus courses spend so much time on continuity before ever reaching FTC.

Section 03

Interactive: FTC Part 1 — Differentiating Accumulation

Pick a function f(t). Watch its accumulation function F(x) get built via integration, then differentiated right back to f(x) — confirming Part 1 in each case.

  FTC Part 1 Explorer
Starting function: f(t) = t
Step 1 — Integrate: F(x) = ∫₀ˣ t dt = x²/2
Step 2 — Differentiate F(x): F'(x) = d/dx [x²/2] = x
F'(x) = x = f(x) ✓ — exactly matches the original function, confirming FTC Part 1
Starting function: f(t) = 2t + 3
Step 1 — Integrate: F(x) = ∫₀ˣ (2t+3) dt = x² + 3x
Step 2 — Differentiate F(x): F'(x) = d/dx [x²+3x] = 2x + 3
F'(x) = 2x + 3 = f(x) ✓ — the accumulation function’s rate of growth is precisely f itself
Starting function: f(t) = cos(t)
Step 1 — Integrate: F(x) = ∫₀ˣ cos(t) dt = sin(x)
Step 2 — Differentiate F(x): F'(x) = d/dx [sin(x)] = cos(x)
F'(x) = cos(x) = f(x) ✓ — works identically even for trigonometric functions

Section 04

Interactive: FTC Part 2 — Evaluating Definite Integrals

Four definite integrals, each solved the FTC way: find an antiderivative, then subtract F(a) from F(b). No rectangles, no limits of Riemann sums.

  FTC Part 2 Explorer
Evaluate: ∫₀³ x² dx
Antiderivative: F(x) = x³/3
F(3) − F(0) = 27/30 = 9 − 0
Result = 9 — this is the exact area under y = x² from x=0 to x=3
Evaluate: ∫₁⁴ (2x+1) dx
Antiderivative: F(x) = x² + x
F(4) − F(1) = (16+4) − (1+1) = 20 − 2
Result = 18 — matches the trapezoid area you’d get geometrically, confirming the shortcut
Evaluate: ∫₀^π sin(x) dx
Antiderivative: F(x) = −cos(x)
F(π) − F(0) = (−cos π) − (−cos 0) = 1 − (−1)
Result = 2 — the area under one full hump of the sine curve
Evaluate: ∫₀² eˣ dx
Antiderivative: F(x) = eˣ
F(2) − F(0) = 1 ≈ 7.389 − 1
Result ≈ 6.389 — exponential functions are their own antiderivative, so this is especially quick

Section 05

The Odometer and Speedometer Model

Every car dashboard already contains a working demonstration of FTC. The speedometer shows your instantaneous speed — a rate. The odometer shows your total distance traveled — an accumulation of that rate over time.

1

The odometer IS an accumulation function

odometer(t) = ∫₀ᵗ speed(τ) dτ + odometer(0). It’s literally integrating your speed over time to build up total distance.

2

The speedometer IS the derivative of the odometer

speed(t) = d/dt [odometer(t)]. At any instant, your speed is exactly how fast the odometer reading is changing. This is FTC Part 1, playing out on your dashboard.

3

Total distance = the odometer difference

Distance traveled between two times a and b = odometer(b) − odometer(a) = ∫ₐᵇ speed(t) dt. This is FTC Part 2 — you don’t need to track every instant of speed, just two odometer readings.

🚗 Why This Analogy Is Exact, Not Approximate

This isn’t a loose metaphor — it’s mathematically precise. Speed genuinely is the derivative of position, and position genuinely is the integral of speed. FTC isn’t describing something *like* a speedometer and odometer; a speedometer and odometer literally are a real-time physical instance of the theorem.

Section 06

Interactive: Watching Area Accumulate

Let f(t) = t (a straight line through the origin). As x grows, the shaded area under the line from 0 to x grows too — that shaded area IS F(x). Click through increasing values of x and watch the accumulation function build up, one snapshot at a time.

  Area Accumulation — f(t) = t
x=1 1 0
F(1) = 1²/2 = 0.5 Shaded area = 0.5
x=2 2 0
F(2) = 2²/2 = 2 Shaded area = 2
x=3 3 0
F(3) = 3²/2 = 4.5 Shaded area = 4.5
x=4 4 0
F(4) = 4²/2 = 8 Shaded area = 8

Notice the area doesn’t grow at a constant rate — it grows FASTER as x increases, because the height (f(x)=x) is also growing. That growing rate of area is exactly f(x) itself: FTC Part 1 in action.

Section 07

Why the Two Parts Are Really One Theorem

Part 1 and Part 2 can feel like two separate facts, but they’re two views of the exact same relationship, just pointed in opposite directions.

Part 1 direction
Start with a function → integrate it → differentiate the result → land back where you started. Integrate, then differentiate = identity.
Part 2 direction
Start with a function → find its antiderivative (undo differentiation) → use it to measure total accumulation. Differentiate, then integrate = identity (up to a constant).

🔄 The Inverse-Operation Pattern

This mirrors √(x²) = x and (√x)² = x — squaring and square-rooting undo each other in both directions. Differentiation and integration have exactly this same inverse relationship, just operating on functions instead of numbers.

Section 08

Worked Example — Projectile Motion

A ball is thrown upward. Its velocity (accounting for gravity) is v(t) = −9.8t + 20 meters per second, where t is measured in seconds. Find the ball’s total displacement between t = 0 and t = 2 seconds.

1

Recognize this as FTC Part 2

Displacement is the integral of velocity. We need ∫₀² v(t) dt, and FTC Part 2 tells us to find an antiderivative and subtract endpoint values.

2

Find the antiderivative

F(t) = −4.9t² + 20t  (check: F'(t) = −9.8t + 20 = v(t) ✓)

3

Evaluate F(2) and F(0)

F(2) = −4.9(4) + 20(2) = −19.6 + 40 = 20.4    F(0) = 0

4

Subtract

Displacement = F(2) − F(0) = 20.4 − 0 = 20.4 meters

v(t) = −9.8t + 20   (velocity)
F(t) = −4.9t² + 20t   (antiderivative — position)

∫₀² v(t) dt = F(2) − F(0)
= (−4.9×4 + 40) − 0
= 20.4

The ball travels 20.4 meters (net) in the first 2 seconds.

⚠ Displacement vs Total Distance

This calculation gives net displacement. If the ball goes up and then starts falling back down within those 2 seconds, the total distance traveled (odometer-style) would be larger, since up-then-down distances partially cancel in a plain integral. Total distance requires integrating |v(t)| instead.

Section 09

Common Mistakes to Avoid

➖

Reversing F(a) and F(b)

The formula is F(b) − F(a), upper bound minus lower bound — never the other way around. Flipping it flips the sign of your answer.

➕

Worrying about “+C”

For definite integrals, the constant of integration always cancels: (F(b)+C) − (F(a)+C) = F(b) − F(a). Any antiderivative works — you don’t need to find “the” one.

🚫

Ignoring discontinuities

If f has a vertical asymptote or jump inside [a,b], you can’t blindly apply FTC across it — the interval must be split or treated as an improper integral.

🔀

Confusing definite and indefinite

∫f(x)dx (indefinite) is a family of functions plus C. ∫ₐᵇf(x)dx (definite) is a single number. Only definite integrals get evaluated via FTC Part 2 directly.

Section 10

Real-World Applications

🚀

Physics — Motion

Displacement from velocity, velocity from acceleration, work done from a variable force — all direct FTC applications in mechanics.

💰

Economics

Total cost is the integral of marginal cost. Consumer and producer surplus are computed as areas — definite integrals evaluated via FTC.

💊

Medicine

Total drug absorbed into the bloodstream over time is the integral of the concentration-rate function — critical for dosage modeling.

⚡

Electrical Engineering

Total electric charge is the integral of current over time: Q = ∫ I(t) dt — a direct real-world use of FTC Part 2.

📊

Probability

A cumulative distribution function (CDF) is the integral of a probability density function (PDF) — FTC Part 1 connects the two directly.

🌱

Biology

Total population growth over a period is the integral of the growth-rate function, used constantly in ecological and epidemiological modeling.

Section 11

Differentiation vs Integration — Comparison Table

AspectDifferentiationIntegration
Geometric meaningSlope of the tangent lineArea under the curve
Physical meaningInstantaneous rate of changeTotal accumulation over an interval
Notationf'(x) or dy/dx∫f(x)dx
Result typeAnother functionA function (indefinite) or a number (definite)
UndoesIntegrationDifferentiation
Everyday exampleSpeedometer readingOdometer reading

Section 12

Code Examples — Python

Verifying FTC symbolically with SymPy

Python
from sympy import *

x, t = symbols('x t')

# FTC Part 1: differentiate an accumulation function
f = t
F = integrate(f, (t, 0, x))       # F(x) = x**2/2
check = diff(F, x)                    # should equal f
print(F, "->", check)              # x**2/2 -> x

# FTC Part 2: evaluate a definite integral
result = integrate(x**2, (x, 0, 3))
print(result)                        # 9

result2 = integrate(sin(x), (x, 0, pi))
print(result2)                       # 2

Numerical integration with SciPy (matching the analytic answer)

Python
from scipy.integrate import quad
import numpy as np

# ∫₀² e^x dx  — should be ~6.389
result, error = quad(lambda x: np.exp(x), 0, 2)
print(f"{result:.4f}")                 # 6.3891

# Projectile displacement: v(t) = -9.8t + 20, from t=0 to t=2
displacement, error = quad(lambda t: -9.8*t + 20, 0, 2)
print(f"{displacement:.1f} meters")     # 20.4 meters

Section 13

Knowledge Quiz

Click a question to expand it, then pick your answer.

FTC establishes that differentiation and integration undo each other — much like addition and subtraction. This is why calculus works as one unified subject, and it’s what makes Part 2’s shortcut for evaluating definite integrals possible.
The speedometer shows instantaneous speed — the rate of change (derivative) of position. The odometer accumulates that speed over time — the integral. speed(t) = d/dt[odometer(t)] is FTC Part 1 in action on your dashboard.
Antiderivative: F(x) = x² + x. F(4) = 16+4 = 20. F(1) = 1+1 = 2. F(4) − F(1) = 20 − 2 = 18.
(F(b)+C) − (F(a)+C) = F(b) − F(a) + C − C = F(b) − F(a). The constant cancels algebraically regardless of its value, which is why you can pick ANY antiderivative when applying FTC Part 2.
Integrating a rate function over an interval gives the TOTAL accumulated quantity over that interval — exactly like integrating speed gives total distance. Here, integrating the fill rate gives total volume added in those 10 minutes.
FTC requires continuity on the entire interval [a,b]. Since 1/x² blows up to infinity at x=0, and 0 is inside [-1,1], the function isn’t continuous across the whole interval — this must be treated as an improper integral with special care, not a direct FTC application.

The post The Fundamental Theorem of Calculus — Explained with Real Examples appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/fundamental-theorem-of-calculus/feed/ 0 804
Real-Life Examples of Permutations and Combinations https://learnwithexamples.org/real-life-examples-of-permutations-and-combinations/ https://learnwithexamples.org/real-life-examples-of-permutations-and-combinations/#respond Mon, 27 Jul 2026 16:15:00 +0000 https://learnwithexamples.org/?p=800 Real-Life Examples of Permutations and Combinations — Explained Simply Learn · With · Examples Lottery odds, poker hands, phone lock codes, Olympic podiums, pizza orders — they’re all counting problems…

The post Real-Life Examples of Permutations and Combinations appeared first on Learn With Examples.

]]>
Real-Life Examples of Permutations and Combinations — Explained Simply

Learn · With · Examples

Lottery odds, poker hands, phone lock codes, Olympic podiums, pizza orders — they’re all counting problems in disguise, and they all come down to one question: does order matter? This guide explains permutations and combinations using nothing but real, everyday examples, with interactive scenario galleries and full worked calculations.

📖 19 min read 🔀 Perm-or-combo identifier 🏅 Permutation gallery 🎟 Combination gallery ❓ 6-question quiz

Section 01

The Difference — A Lock Code vs. A Fruit Bowl

Imagine two everyday situations. First: you’re setting a 3-digit code on a padlock using digits 1, 2, and 3. The code 1-2-3 is completely different from 3-2-1 — they open different locks (or rather, only the exact sequence you set will open yours). Order matters. This is a permutation.

Second: you’re picking 3 fruits from a bowl to make a smoothie — an apple, a banana, and an orange. It doesn’t matter if you grabbed the apple first or the orange first — you end up with the exact same smoothie ingredients either way. Order doesn’t matter. This is a combination.

💡 The One Question That Decides Everything

Every counting problem in this entire topic reduces to a single question: if I rearrange the same items, do I get something different? Yes → permutation. No → combination. Everything else is just formula mechanics built on top of that one distinction.

n!
Factorial — the building block behind every permutation and combination formula
nPr
Permutations — ordered selections
nCr
Combinations — unordered selections
1/13.98M
Odds of matching all 6 numbers in a 6/49 lottery — a real combination calculation

Section 02

Formal Definitions — The nPr and nCr Formulas

Both formulas start from the same n items, choosing r of them — they only differ in whether order counts.

P(n,r) = n! / (n − r)!
Permutations — the number of ORDERED ways to choose r items from n
C(n,r) = n! / [r! (n − r)!]
Combinations — the number of UNORDERED ways to choose r items from n

🔗 How They’re Related

Notice combinations = permutations ÷ r!. That’s not a coincidence — every unordered group of r items can be arranged in r! different orders. Combinations “collapse” all those orderings into one, because order doesn’t matter. C(n,r) = P(n,r) / r!

Section 03

Interactive: Permutation or Combination?

Click through 5 real scenarios. Try to guess before reading the reasoning — this is the exact judgment call you’ll need to make on every word problem you encounter.

  Scenario Identifier
“Assign gold, silver, and bronze medals to 3 of 10 finalists.”
PERMUTATION
Gold, silver, and bronze are three different roles. Alice-gold/Bob-silver is a completely different outcome from Bob-gold/Alice-silver, even though the same two people are involved. Order (which position each person lands in) changes the result.
“Choose 4 pizza toppings from a menu of 12.”
COMBINATION
A pizza with pepperoni-mushroom-olive-onion is the exact same pizza as onion-olive-mushroom-pepperoni. The order you listed the toppings in doesn’t create a different pizza — only the final set of 4 toppings matters.
“Set a 4-digit lock code using digits 0–9.”
PERMUTATION
1-2-3-4 and 4-3-2-1 are different codes that open different locks. Position matters — this is a permutation (specifically one that allows repeated digits, covered later).
“Pick 2 co-captains from a 15-person sports team.”
COMBINATION
Both co-captains hold the identical role — there’s no “first co-captain” and “second co-captain.” Choosing Sam-then-Priya gives the exact same pair of co-captains as Priya-then-Sam.
“Arrange 6 unique paintings in a row along a gallery wall.”
PERMUTATION
Each painting occupies a specific position on the wall (1st, 2nd, 3rd…). Swapping two paintings’ positions creates a visibly different arrangement — position matters, so this is a permutation.

Section 04

The Factorial Foundation

Both formulas are built entirely from factorials, so it’s worth getting comfortable with them first. n! (read “n factorial”) means multiply every whole number from n down to 1.

5! = 5 × 4 × 3 × 2 × 1 = 120
The number of ways to arrange 5 distinct items in a row
nn!Real meaning
0!1By definition — there’s exactly 1 way to arrange nothing
1!1One item — only one arrangement possible
2!22 books on a shelf: AB or BA
3!63 runners crossing the finish line, in order
4!244 people seated around a table
5!1205 songs in a playlist order
6!7206 books arranged on a shelf
10!3,628,80010 runners in a full race ranking

⚠ Why 0! = 1

This trips people up constantly. Think of factorial as “number of ways to arrange these items.” With zero items, there’s exactly one way to arrange them: do nothing. An empty arrangement is still one valid arrangement — so 0! = 1, not 0.

Section 07

The “Does Order Matter?” Test

A reliable 3-step process for any word problem you encounter:

1

Swap two of your chosen items — does the outcome change?

If choosing Alice-then-Bob gives a genuinely different result than Bob-then-Alice (different roles, different positions), order matters → permutation. If it’s the same outcome either way → combination.

2

Check for distinct roles or positions

Ranks (1st/2nd/3rd), seats, digit positions, and named roles (president/treasurer) all signal permutation. A “set” or “group” with no internal distinction signals combination.

3

Check if repetition is allowed

If the same item can be chosen more than once (like PIN digits or dice rolls), you need the “with repetition” formula — covered in Section 9 — regardless of whether order matters.

⚠ The Most Common Mistake

Students often default to combinations because the formula “feels simpler.” But most everyday scenarios with roles, rankings, sequences, or codes are permutations. Always run the swap test in Step 1 before picking a formula.

Section 08

Worked Example — Building a Real Password System

Let’s combine everything into one realistic problem: a website requires passwords with exactly 2 different letters (no repeats) followed by 3 digits (digits CAN repeat). How many total passwords are possible?

1

Break it into two independent parts

Part A: choosing 2 different letters in order. Part B: choosing 3 digits, repeats allowed. We’ll solve each separately, then multiply.

2

Solve Part A — the letters

2 letters, no repeats, order matters (AB ≠ BA as passwords) → this is a permutation. P(26,2) = 26 × 25 = 650 possible letter pairs.

3

Solve Part B — the digits

3 digits, repeats allowed, order matters → permutation with repetition. 10 × 10 × 10 = 1,000 possible digit sequences.

4

Apply the multiplication principle

Every letter-pair can be combined with every digit-sequence — so multiply the two results together.

Part A (letters): P(26,2) = 26 × 25 = 650
Part B (digits): 10³ = 1,000

Total passwords = 650 × 1,000
= 650,000 possible passwords

🔗 The Multiplication Principle

Whenever a problem splits into independent stages (letters, THEN digits), multiply the number of possibilities at each stage. This single rule — the fundamental counting principle — is what lets you break any complex real-world counting problem into smaller, solvable permutation and combination pieces.

Section 09

Permutations and Combinations with Repetition

The formulas above assume each item can only be chosen once. Real life often allows repeats — here’s how the math changes.

Permutations WITH repetition — n choices, repeated r times, order matters. This is how PIN codes and license plates are counted.
C(n + r − 1, r)
Combinations WITH repetition (“stars and bars”) — choosing r items from n types, repeats allowed, order doesn’t matter

Worked Example — Ice Cream Scoops

An ice cream shop has 5 flavors. You order 3 scoops, and you’re allowed to repeat a flavor (like 2 scoops chocolate + 1 scoop vanilla). Order doesn’t matter — a cup with chocolate-chocolate-vanilla is the same order regardless of which scoop went in first.

n = 5 flavors, r = 3 scoops, repetition allowed, order doesn’t matter
C(n + r − 1, r) = C(5 + 3 − 1, 3) = C(7,3)
= 7! / [3! × 4!] = (7×6×5) / 6
= 35 possible ice cream orders

Section 10

Pascal’s Triangle Connection

Every entry in Pascal’s Triangle IS a combination value. Row n, position k gives you exactly C(n,k) — no calculation required, just count and look it up.

n=01
n=11   1
n=21   2   1
n=31   3   3   1
n=41   4   6   4   1
n=51   5   10   10   5   1
n=61   6   15   20   15   6   1

Look at row n=6: the entries are C(6,0)=1, C(6,1)=6, C(6,2)=15, C(6,3)=20, and so on. Each number is the sum of the two numbers diagonally above it — which is exactly Pascal’s Rule: C(n,k) = C(n−1,k−1) + C(n−1,k).

💡 Where You’ve Actually Seen This

Pascal’s Triangle is also the coefficients of a binomial expansion — (a+b)⁶ expands using exactly the row-6 numbers: 1, 6, 15, 20, 15, 6, 1. It’s the same combinatorics showing up in algebra, probability, and even genetics (Punnett square ratios follow these same patterns).

Section 11

Real-World Domains

🔐

Cryptography

Key space size for a cipher is a permutation calculation — the number of possible keys directly determines how brute-forceable an encryption scheme is.

🧬

Genetics

Counting possible DNA sequences, or the number of ways alleles can combine in offspring, relies directly on permutation and combination formulas.

🏆

Tournament Brackets

The number of possible ways a single-elimination bracket can play out, or how many unique round-robin schedules exist, is pure combinatorics.

📅

Scheduling

Assigning employees to shifts, or students to exam time slots, is a permutation/combination problem — especially when constraints (no repeats, fixed roles) apply.

🧪

Quality Control

Choosing a random sample of r items from a batch of n to inspect for defects is a classic combination — order of inspection doesn’t matter.

🎲

Game Design

Loot drop tables, card game deck compositions, and probability balancing in games all rely on combinatorics to calculate fair odds.

Section 12

Code Examples — Python

Using math.perm() and math.comb() (Python 3.8+)

Python
import math

# Olympic podium — 3 medals from 8 finalists, order matters
print(math.perm(8, 3))     # 336

# Lottery — choose 6 numbers from 49, order doesn't matter
print(math.comb(49, 6))    # 13983816

# Password letters — 2 of 26, no repeats, order matters
print(math.perm(26, 2))    # 650

# Pizza toppings — choose 3 of 8, order doesn't matter
print(math.comb(8, 3))     # 56

# PIN code — permutation WITH repetition (not built into math module)
pin_combos = 10 ** 4
print(pin_combos)          # 10000

# Ice cream scoops — combination WITH repetition (stars and bars)
ice_cream = math.comb(5 + 3 - 1, 3)
print(ice_cream)          # 35

Generating the actual arrangements with itertools

Python
from itertools import permutations, combinations

runners = ['A', 'B', 'C', 'D']

# All ordered podium outcomes (top 3 of 4 runners)
podiums = list(permutations(runners, 3))
print(len(podiums))      # 24  =  4P3
print(podiums[:3])
# [('A','B','C'), ('A','B','D'), ('A','C','B'), ...]

# All unordered committees (any 3 of 4 people)
committees = list(combinations(runners, 3))
print(len(committees))   # 4  =  4C3
print(committees)
# [('A','B','C'), ('A','B','D'), ('A','C','D'), ('B','C','D')]

Section 13

Knowledge Quiz

Click a question to expand it, then pick your answer.

If swapping the order of your chosen items changes the result (different roles, positions, or sequence), it’s a permutation. If the same items in any order count as the same outcome, it’s a combination.
0! = 1 by definition. Think of factorial as “number of ways to arrange these items” — with zero items, there’s exactly one way to arrange them: the empty arrangement. This convention also keeps the nPr and nCr formulas working correctly when r = n.
The three roles are distinct (line leader ≠ door holder ≠ paper collector), so order/role matters — this is a permutation. P(20,3) = 20 × 19 × 18 = 6,840.
A hand of {A♠, K♥, 7♦, 3♣, 2♠} is the same hand no matter which card was dealt first. Since order doesn’t affect the outcome, it’s counted with combinations: C(52,5) = 2,598,960.
PIN digits can repeat (like 1-1-2-2) and order matters (position 1 ≠ position 2). This is “permutation with repetition”: n choices raised to the power of r positions → 10⁴ = 10,000.
Every entry in Pascal’s Triangle is a combination: row n, position k gives C(n,k). Row 6’s entries (1,6,15,20,15,6,1) are C(6,0) through C(6,6) — and indeed C(6,3) = 20.

The post Real-Life Examples of Permutations and Combinations appeared first on Learn With Examples.

]]>
https://learnwithexamples.org/real-life-examples-of-permutations-and-combinations/feed/ 0 800