Lossless compression makes a file smaller and lets you rebuild the exact original. Lossy compression makes it smaller still by permanently removing some data, so the original cannot be recovered. Questions ask you to define each, choose one for a situation and justify the choice.
This lesson follows from sampled audio size because compression exists to reduce the large sizes you have just calculated.
What is the difference, in plain terms?
Lossless methods find patterns and store them more efficiently. Nothing is thrown away. Text files, program files and data that must stay exact need this.
Lossy methods discard detail that people are unlikely to notice. Photos in JPEG form and audio in MP3 form are everyday examples. Quality drops a little, and the more you compress, the more is lost.
| Feature | Lossless | Lossy |
|---|---|---|
| Original recoverable | Yes | No |
| File size reduction | Smaller | Usually much smaller |
| Quality | Unchanged | Reduced |
| Suits | Text, programs, medical scans | Photos, music, video for viewing |
How does run-length encoding work?
Run-length encoding (RLE) is a lossless method. It replaces a run of the same value with the value and a count.
The idea as Cambridge-style pseudocode, for a string stored in Data:
Count ← 1
FOR i ← 2 TO LENGTH(Data)
IF MID(Data, i, 1) = MID(Data, i - 1, 1) THEN
Count ← Count + 1
ELSE
OUTPUT MID(Data, i - 1, 1), Count
Count ← 1
ENDIF
NEXT i
OUTPUT MID(Data, LENGTH(Data), 1), Count
Trace with Data = “AAABCC” (6 characters):
| i | Compared | Action | Count |
|---|---|---|---|
| start | 1 | ||
| 2 | A, A | same, add 1 | 2 |
| 3 | A, A | same, add 1 | 3 |
| 4 | B, A | differ: output A 3 | 1 |
| 5 | C, B | differ: output B 1 | 1 |
| 6 | C, C | same, add 1 | 2 |
| after loop | output C 2 |
Result: A3 B1 C2. You can step through a loop like this in the pseudocode trace trainer.
Worked example
One row of a black-and-white image is stored as W W W W W W W W B B W W W W (14 pixels, W = white, B = black). Compress it with RLE and say whether the method is lossless.
Step 1, group the runs: 8 whites, 2 blacks, 4 whites.
Step 2, write value and count: W8 B2 W4.
Step 3, check by decoding: W×8, B×2, W×4 gives the original 14 pixels exactly.
Conclusion: 14 values became 3 pairs, which is 6 stored items. Because decoding returns the original exactly, RLE is lossless.
The mistake to watch for
A common slip is to assume compression always shrinks the data.
Mistaken answer: “ABCD” compresses to A1B1C1D1, so it is smaller.
The student counted characters wrongly. A1B1C1D1 has 8 characters, more than the original 4.
The correction: RLE only helps when values repeat. Data with no runs can grow. A good answer says that the saving depends on the pattern in the data.
Check yourself
Try these, then open each answer.
1. Encode WWWBBWWWWW with run-length encoding.
Show answer
Runs: 3 W, 2 B, 5 W. Answer: W3 B2 W5. Decoding gives WWW BB WWWWW, the original 10 characters.
2. A hospital stores a scan used for diagnosis. Should it use lossless or lossy compression? Give a reason.
Show answer
Lossless. Lossy compression permanently removes detail, and missing detail in a medical image could change a diagnosis. Lossless keeps every pixel.
3. Name one reason a streaming music service might choose lossy compression.
Show answer
Lossy files are much smaller, so they download faster and use less data, and the loss of quality is hard for most listeners to notice.
Where this leads next
Choosing between methods is really a trade-off between size and quality, so continue with explaining a quality-storage trade-off. Then use the practice set for mixed questions.
If definitions are secure but the scenario answers feel loose, a teacher can drill them in online one-to-one Computer Science tuition.