VISION
คู่มือ · อ่านเองทีละบทHandbook · read it on your own

05 · Neural networks · โครงข่ายประสาทเทียม

อ่านเองได้ทีละบท ทุกบทมาพร้อมแบบฝึกหัดให้ลองทำเองและคำถามท้ายบท เมื่ออยากรู้ให้ลึกกว่านี้ กลับไปที่ห้องเรียนบนเว็บเพื่อทดลองของจริง

Read a chapter at a time. Each comes with an exercise you can try on your own and questions to check yourself. When you want to go deeper, go back to the room on this site and run it for real.

ในหน้านี้On this page
  1. 1. From chosen numbers to learned numbers
  2. 2. Why "deep": layers build on layers
  3. 3. The network on this site: MobileNetV2
  4. 4. How the detector reuses the same network
  5. 5. What "the network knows" — and does not
  6. 6. A short history (detailed in the reading list)
  7. Try it yourself · ลองทำเอง
  8. Check yourself
In one sentence: a neural network is many layers of convolutions whose numbers were found by showing it millions of labelled pictures and nudging every number, a little at a time, towards fewer mistakes.

← 04 · Handbook · Site: /learn, chapter 7 · Code: public/js/ml/


1. From chosen numbers to learned numbers

Chapters 01–04 used rules a person wrote: a brightness formula, a threshold, Sobel's nine numbers, a motion delta. They are transparent and they fail on anything their author did not anticipate. Nobody can write the kernel for "motorcycle helmet seen from behind at dusk".

The change that reshaped the field around 2012 was to keep the structure — sliding windows, stacked — and let the numbers be learned:

  1. Start with random numbers in every kernel.
  2. Show the network a labelled picture ("this is a car"). It outputs a score for every class — a class being one of the categories it may choose between (a car, a bus, a person; ImageNet has 1,000).
  3. Measure how wrong it was (the loss).
  4. Work out, for every one of its millions of numbers, which direction would have made it slightly less wrong (backpropagation computes this; Rumelhart, Hinton & Williams, 1986).
  5. Move every number a tiny step in that direction (gradient descent).
  6. Repeat for millions of pictures, many times over.

Chapter 07 does steps 2–6 in front of you, on one layer, in plain JavaScript — the loss curve on /train is this process.

2. Why "deep": layers build on layers

Each layer convolves the previous layer's output. The first layer sees pixels; the second sees "where the first layer's edge detectors fired"; and so on. When researchers visualised what trained layers respond to (Zeiler & Fergus, 2014), a hierarchy appeared:

DepthResponds toExample
layer 1edges, colour blobsrediscovers Sobel-like filters (chapter 03)
layers 2–3corners, textures, simple patternsmesh, stripes, tread
middlepartswheels, eyes, windows
lateobject-like wholescar fronts, faces, dog heads

Nobody programmed this hierarchy. It emerged because it is the most useful way to turn pixels into labels.

3. The network on this site: MobileNetV2

The site runs MobileNetV2 (Sandler et al., 2018) — designed at Google to be accurate enough while running on a phone.

MobileNetV2: smaller in space, deeper in meaning

Reading the shapes left to right (from the paper's Table 2): the 224 × 224 × 3 input is repeatedly halved in width and height while the number of channels grows — less "where", more "what" — until a 7 × 7 grid with 1,280 channels is averaged into 1,280 numbers. Those 1,280 numbers are the network's summary of the whole picture: its embedding, or "fingerprint". A final layer turns them into scores for ImageNet's 1,000 classes.

MobileNetV2 1.0 · 224
Weights3.4 million
Arithmetic per picture~300 million multiply-adds
Trained onImageNet: 1.28 million photos, 1,000 classes (Deng et al., 2009)
ImageNet top-1 accuracy (paper)72.0%
On this site14 MB of weights, self-hosted at /models/mobilenet_v2_1.0_224/
Used forthe 1,280-number fingerprint on /train (chapter 07)

The site uses only the fingerprint, not the ImageNet guesses. public/js/ml/embedder.js can return the top guesses (look(frame, { topk: 3 })), but ImageNet's 1,000 English labels — more than a hundred of them dog breeds — say little useful about a Thai road, and we will not machine-translate them.

The trick that makes it small: depthwise separable convolution

A normal convolution layer mixes space and channels at once: every output channel looks at a 3×3 window across all input channels. MobileNets (Howard et al., 2017) split this into two cheap steps:

  1. Depthwise: one 3×3 kernel per input channel (space only).
  2. Pointwise: a 1×1 convolution mixing channels (channels only).

For a 3×3 kernel this costs roughly 8–9× less arithmetic for a small loss in accuracy. MobileNetV2 adds inverted residuals (expand channels ×6, filter, compress, and add the input back) and linear bottlenecks (no ReLU — the clamp-at-zero activation — on the compressed layer, which would destroy the information). That is how 3.4 million weights get to 72% on ImageNet — AlexNet (2012) used about 60 million to reach a lower accuracy.

4. How the detector reuses the same network

Object detection (chapter 06) uses MobileNetV2 again, as a backbone: the same layers produce feature grids at several scales, and small extra layers (the SSDLite "head") predict boxes and classes from them. One feature extractor, two jobs — the reason one family of networks serves both /learn and /train.

5. What "the network knows" — and does not

6. A short history (detailed in the reading list)

YearMoment
1959Hubel & Wiesel find edge-sensitive cells in the cat visual cortex
1966MIT's "Summer Vision Project" expects to solve vision in a summer
1980Fukushima's Neocognitron: layered, convolution-like, self-organising
1986Backpropagation popularised (Rumelhart, Hinton & Williams)
1998LeCun et al.: convolutional networks read handwritten cheques
2009ImageNet: 14 million labelled images (Deng et al.); MobileNetV2 was trained on the 1,000-class, 1.28 million slice of them
2012AlexNet wins ImageNet: top-5 error 15.3% vs 26.2% for the runner-up
2015ResNet: 152 layers via residual connections (He et al.)
2017–18MobileNets, MobileNetV2: deep vision fits on a phone
2020Vision Transformers: attention instead of convolution (Dosovitskiy et al.)
2021CLIP: learning from images paired with text (Radford et al.)
2023Segment Anything: one model that outlines anything (Kirillov et al.)

Try it yourself · ลองทำเอง

เป้าหมาย · Goal: ดูเครื่องมีความมั่นใจและเสียความมั่นใจอย่างไร เมื่อภาพเปลี่ยนเพียงเล็กน้อย / Watch a network gain and lose confidence when the picture changes only a little.

ขั้นตอน · Steps

  1. เปิด /learn บท 7 กด "โหลดโครงข่ายประสาทเทียม" โหลดครั้งเดียว ประมาณ 14 MB แล้วทำงานในเบราว์เซอร์ของคุณ — Open /learn chapter 7 and press Load the neural network: about 14 MB, once, and it runs in your browser.
  2. ชี้ไปที่วัตถุสามอย่าง จดชื่อและตัวเลขความมั่นใจแต่ละอย่าง — Point at three objects and write down each label and its score.
  3. หาวัตถุหนึ่งที่เครื่องไม่เห็นเลย แล้วขยับกล้องเข้าใกล้ทีละนิด ดูคะแนนเปลี่ยนแค่ไหน — Find one object it misses, then move a little closer and watch the score move.
  4. เลื่อน "แสดงเมื่อมั่นใจอย่างน้อย" จาก 0 ขึ้นเป็น 50% แล้วดูกรอบที่เหลือ — Raise Show only when at least this sure from 0 to 50% and see which boxes survive.

ควรเห็น · You should see

ถ้าไม่เห็น · If you do not — ถ้าไม่มีกรอบเลย ให้ตรวจว่ากดโหลดแล้ว และลดเกณฑ์ความมั่นใจลง — If no box ever appears: check that the network loaded, and lower the confidence bar.

Check yourself

1. Why does the picture get smaller in width and height but deeper in channels as it passes through MobileNetV2?

Each step trades location detail for meaning. Early on, knowing exactly where an edge is matters; by the end, the network needs to know what is present more than precisely where, and it needs many channels to represent many kinds of "what".

2. A classifier trained on dogs and cats is shown a car. What does it output?

Scores for "dog" and "cat" that sum to 1 — perhaps "cat 0.71". It has no way to say "neither" unless it was given such a class or a separate mechanism for detecting unfamiliar inputs.

3. Why can a 3.4-million-weight network beat AlexNet's ~60 million?

Better architecture: depthwise separable convolutions spend arithmetic where it matters, residual connections let deeper networks train, and years of improved training methods. Size is not the same as capability.


สรุปภาษาไทย

โครงข่ายประสาทเทียม คือคอนโวลูชันหลายชั้นซ้อนกัน แต่ตัวเลขในทุกหน้าต่างไม่ได้มาจากคนเลือก มาจากการ ฝึก: เริ่มด้วยตัวเลขสุ่ม ให้ดูภาพที่มีป้ายกำกับ วัดว่าผิดแค่ไหน (loss) แล้วขยับตัวเลขทุกตัวไปในทิศที่ผิดน้อยลงทีละนิด (gradient descent) ทำซ้ำกับภาพนับล้าน

ชั้นต้น ๆ เรียนรู้ขอบและสี ชั้นกลางเรียนรู้พื้นผิวและชิ้นส่วน ชั้นท้ายเรียนรู้รูปร่างที่คล้ายวัตถุ ไม่มีใครโปรแกรมลำดับขั้นนี้ มันเกิดขึ้นเองเพราะเป็นวิธีที่ได้ผลที่สุด

เว็บไซต์ใช้ MobileNetV2 (2018) ซึ่งมีน้ำหนัก 3.4 ล้านตัว ฝึกจากภาพ ImageNet 1.28 ล้านภาพ ภาพขนาด 224×224 ถูกย่อลงทีละครึ่งแต่มีช่องข้อมูลมากขึ้น จนเหลือ ตัวเลข 1,280 ตัว ที่เปรียบเหมือน "ลายนิ้วมือ" ของภาพ ห้อง /train ใช้ลายนิ้วมือนี้ ส่วนเครื่องตรวจจับวัตถุใช้โครงข่ายเดียวกันเป็นฐาน

ข้อควรจำ: โครงข่ายรู้แค่สถิติของภาพที่ใช้ฝึก คะแนนความมั่นใจไม่ใช่ความถูกต้อง การเปลี่ยนพิกเซลเล็กน้อยอาจทำให้คำตอบเปลี่ยน และมันพูดว่า "ไม่รู้" ไม่เป็น เว้นแต่จะถูกออกแบบมาให้พูดได้

Next: 06 · Object detection →