เครื่องมองอย่างไร
ทีละขั้นHow it looks,
step by step
ทุกบทศึกษาภาพเดียวกัน เลือกได้จากแถบสีดำด้านล่าง จะเป็นกล้องสาธารณะ กล้องของคุณเอง หรือรูปของคุณก็ได้ เปลี่ยนเมื่อไรก็ได้ เครื่องมือทุกบทจะเปลี่ยนตาม
Every chapter studies the same picture, chosen in the black bar below: a public camera, your own camera, or a photo of yours. Change it any time and every instrument follows.
ภาพจากกล้องหรือรูปของคุณประมวลผลในเบราว์เซอร์นี้เท่านั้น ไม่มีการส่งหรือบันทึกไว้ที่ไหนYour camera or photo is processed in this browser only. Nothing is sent or saved anywhere.
ภาพคือตัวเลขA picture is numbers
ภาพบนจอประกอบด้วยช่องสี่เหลี่ยมเล็ก ๆ เรียกว่าพิกเซล คอมพิวเตอร์ไม่ได้เห็นถนนหรือรถ มันได้รับแค่ตัวเลขบอกความสว่างของแต่ละช่อง ตั้งแต่ 0 (ดำสนิท) ถึง 255 (ขาวที่สุด)
A picture on a screen is made of tiny squares called pixels. The computer does not see a road or a car. It receives one number for how bright each square is, from 0 (black) to 255 (white).
เลื่อนแถบให้เหลือช่องน้อยลง แล้วตัวเลขจะปรากฏในช่อง ภาพจากกล้องจริงหนึ่งภาพมีตัวเลขแบบนี้หลายล้านตัว ทุกอย่างที่เหลือในหน้านี้คือการคิดเลขกับตารางนี้
Slide to fewer squares and the numbers appear inside them. One real camera picture holds millions of these. Everything else on this page is arithmetic on this grid.
ลองดูTry หาจุดที่สว่างที่สุดในภาพ ตัวเลขของมันใกล้ 255 ไหมFind the brightest spot. Is its number close to 255?
สีคือตัวเลขสามตัวColour is three numbers
ภาพสีเก็บตัวเลขสามตัวต่อหนึ่งพิกเซล คือ แดง เขียว และน้ำเงิน (R G B) แต่ละตัวมีค่า 0 ถึง 255 จอของคุณผสมแสงสามสีนี้เข้าด้วยกัน ตาเราจึงเห็นเป็นสีเดียว
A colour picture stores three numbers per pixel: red, green and blue (R G B), each from 0 to 255. Your screen mixes those three lights, and your eye sees one colour.
ดูทีละช่องสีแล้วจะเห็นว่าท้องฟ้าสว่างในช่องน้ำเงิน ต้นไม้สว่างในช่องเขียว ไฟท้ายรถสว่างในช่องแดง เครื่องใช้ความต่างแบบนี้แยกสิ่งต่าง ๆ ออกจากกันได้ ก่อนจะรู้ด้วยซ้ำว่าสิ่งนั้นคืออะไร
Look at one channel at a time: sky is bright in blue, trees in green, tail-lights in red. A machine can use differences like these to tell things apart before it knows what anything is.
ลองดูTry ถือของสีแดงไว้หน้ากล้องของคุณ แล้วดูว่ามันสว่างในช่องไหนHold something red up to your camera. Which channel lights up?
สว่างหรือมืดLight or dark
วิธีที่ง่ายที่สุดในการแยกสิ่งของออกจากพื้นหลัง คือเลือกตัวเลขหนึ่งตัวเป็นเส้นแบ่ง พิกเซลที่สว่างกว่าเส้นนี้นับเป็น “สว่าง” ที่เหลือนับเป็น “มืด” วิธีนี้เรียกว่าการตัดด้วยค่าขีดแบ่ง (thresholding)
The simplest way to separate something from its background: pick one number as a line. Pixels brighter than the line count as “light”, the rest as “dark”. This is called thresholding.
กราฟใต้ภาพคือฮิสโตแกรม นับว่ามีพิกเซลกี่จุดในแต่ละระดับความสว่าง ปุ่ม “อัตโนมัติ” ใช้วิธีของโอตสึ (Otsu, 1979) หาเส้นแบ่งที่แยกพิกเซลเป็นสองกลุ่มได้ชัดที่สุด เครื่องสแกนเอกสารยังใช้หลักนี้อยู่ทุกวันนี้
The chart under the picture is a histogram: how many pixels sit at each brightness. “Automatic” uses Otsu's method (Otsu, 1979) to find the line that splits the pixels most cleanly into two groups. Document scanners still work this way.
ลองดูTry ส่องกล้องไปที่กระดาษที่มีตัวหนังสือ แล้วกด “อัตโนมัติ”Point your camera at a page of writing and press “Automatic”.
หน้าต่างเล็ก ๆ ที่เลื่อนไปSliding a small window
วางหน้าต่างขนาด 3×3 ช่องลงบนภาพ เอาตัวเลขเก้าตัวใต้หน้าต่างคูณกับตัวเลขเก้าตัวของ “เคอร์เนล” แล้วบวกกันทั้งหมด ได้พิกเซลใหม่หนึ่งจุด จากนั้นเลื่อนไปหนึ่งช่องแล้วทำซ้ำจนทั่วภาพ ขั้นตอนนี้เรียกว่าคอนโวลูชัน (convolution)
Lay a 3×3 window on the picture. Multiply the nine numbers under it by the nine numbers of a “kernel”, and add them all up: that is one new pixel. Slide one step and repeat, across the whole picture. This is called convolution.
เปลี่ยนตัวเลขเก้าตัว ผลก็เปลี่ยน ภาพอาจเบลอ คมขึ้น หรือเหลือแต่ขอบ แตะที่ภาพเพื่อย้ายหน้าต่าง แล้วดูการคำนวณจริงด้านล่าง
Change the nine numbers and the result changes: blur, sharpen, or nothing but edges. Tap the picture to move the window and watch the real arithmetic below it.
นี่คือเมล็ดพันธุ์ของโครงข่ายประสาทเทียมแบบคอนโวลูชัน (CNN) ต่างกันเพียงว่าไม่มีใครเลือกตัวเลขให้ CNN มันเรียนรู้ตัวเลขเองจากภาพตัวอย่าง และซ้อนหน้าต่างแบบนี้หลายพันอันเป็นชั้น ๆ (LeCun และคณะ, 1998; Krizhevsky และคณะ, 2012)
This is the seed of convolutional neural networks (CNNs). The only difference: nobody chooses a CNN's numbers. It learns them from example pictures, and stacks thousands of such windows in layers (LeCun et al., 1998; Krizhevsky et al., 2012).
ขอบEdges
ขอบคือตรงที่ความสว่างเปลี่ยนเร็ว เช่น ตรงที่ตัวรถตัดกับพื้นถนน ตัวกรองโซเบล (Sobel) ใช้หน้าต่างสองแบบจากบทที่แล้ว แบบหนึ่งหาการเปลี่ยนจากซ้ายไปขวา อีกแบบหาจากบนลงล่าง แล้วรวมกันเป็นความแรงของขอบ
An edge is where brightness changes fast, like where a car meets the road. The Sobel filter uses two windows from the last chapter, one for left-to-right change and one for top-to-bottom, and combines them into edge strength.
ยิ่งสว่างยิ่งเป็นขอบที่ชัด เพิ่มความไวเพื่อดูขอบจาง ๆ เครื่องยังไม่รู้ว่าเส้นเหล่านี้คืออะไร แต่รูปร่างเริ่มโผล่ขึ้นมาแล้ว ชั้นแรก ๆ ของ CNN ที่ฝึกเสร็จแล้วมักเรียนรู้ตัวกรองที่หน้าตาคล้ายแบบนี้ขึ้นมาเอง
Brighter means a stronger edge. Raise the gain to see faint ones. The machine still has no idea what any line is, but shape is starting to appear. The first layers of a trained CNN often end up learning filters much like this on their own.
ลองดูTry ใช้กล้องของคุณแล้วขยับมือช้า ๆ ขอบนิ้วชัดแค่ไหนเมื่อพื้นหลังสว่างหรือมืดUse your camera and move a hand slowly. How clear are your fingers against a light or a dark background?
อะไรขยับWhat moved
เอาภาพตอนนี้ลบด้วยภาพเมื่อครู่ ช่องที่ตัวเลขเปลี่ยนมากกว่าค่าที่ตั้งไว้คือสิ่งที่ขยับ สีส้มคือสิ่งที่เครื่องสังเกตเห็น กรอบคือกลุ่มพิกเซลที่อยู่ติดกัน
Subtract the picture from a moment ago from the picture now. Squares whose number changed by more than the setting are things that moved. Orange marks what the machine noticed; boxes group pixels that touch.
ตั้งค่าต่ำเกินไป เครื่องจะจับได้แม้แต่ใบไม้ไหวหรือสัญญาณรบกวนของกล้อง ตั้งสูงเกินไป มันจะพลาดรถที่สีคล้ายถนน ไม่มีค่าไหนถูกต้องเสมอ
Set it too low and it catches swaying leaves and camera noise. Too high and it misses a car the same colour as the road. No setting is always right.
กล้องบางตัวส่งภาพนิ่งทุกราว 20 วินาที ไม่ใช่วิดีโอ เครื่องจึงเห็นความเปลี่ยนแปลงได้เฉพาะตอนที่ภาพใหม่มาถึง และในยี่สิบวินาที รถทั้งคันอาจผ่านมาแล้วก็ผ่านไป
Some cameras send a still picture about every 20 seconds, not video. The machine can only see change when the next picture arrives — and in twenty seconds a whole car can come and go.
ลองดูTry เปิดกล้องของคุณ อยู่นิ่ง ๆ แล้วโบกมือOpen your camera, keep still, then wave.
โครงข่ายที่ฝึกแล้วลองเดาA trained network guesses
ทุกอย่างก่อนหน้านี้ทำตามกฎที่คนเขียนไว้ บทนี้ต่างออกไป โครงข่ายประสาทเทียมชื่อ SSDLite-MobileNetV2 เรียนรู้ว่า “รถยนต์” “คน” หรือ “เรือ” หน้าตาเป็นอย่างไร จากภาพที่มีคนติดป้ายไว้กว่าแสนภาพในชุดข้อมูล COCO (Lin และคณะ, 2014) มันดูภาพเพียงครั้งเดียว แล้วเสนอกรอบหลายพันกรอบ แต่ละกรอบมีคะแนน
Everything so far followed rules a person wrote. This is different. A neural network called SSDLite-MobileNetV2 learned what “car”, “person” and “boat” look like from over a hundred thousand labelled photos in the COCO dataset (Lin et al., 2014). It looks at the picture once and proposes thousands of boxes, each with a score.
คะแนนคือความมั่นใจของตัวมันเอง ไม่ใช่ความจริง เลื่อนแถบเพื่อเลือกว่ามันต้องมั่นใจแค่ไหนถึงจะพูด ตั้งต่ำไปจะเห็นของที่ไม่มีอยู่จริง ตั้งสูงไปจะพลาดของที่มีอยู่ ทุกระบบจริงต้องเลือกจุดแลกเปลี่ยนนี้
The score is its own confidence, not the truth. Slide to choose how sure it must be before it speaks. Too low and it sees things that are not there; too high and it misses things that are. Every real system has to choose this trade-off.
โมเดลมีขนาด 18 MB โหลดครั้งเดียวเมื่อคุณกดปุ่ม และคำนวณในเบราว์เซอร์นี้ ไม่มีภาพถูกส่งออกไปไหน
The model is 18 MB, loaded once when you press the button, and computes in this browser. No picture is sent anywhere.
มันรู้ได้อย่างไรว่านี่คือสุนัข แมว หรือรถHow does it know: dog, cat or car?
ไม่มีใครบอกมันว่า “สุนัขมีหูตก” มันดูภาพที่มีป้ายกำกับนับแสนภาพ แล้วปรับตัวเลขข้างในทีละนิด จนชิ้นส่วนที่มักอยู่ในภาพสุนัขช่วยดันคะแนน “สุนัข” ขึ้น
Nobody told it “a dog has floppy ears”. It looked at hundreds of thousands of labelled photos and nudged its numbers, little by little, until the parts that tend to appear in dog photos push the “dog” score up.
สิ่งที่มันทำไม่ได้What it cannot do
เลื่อนแถบเพื่อลดจำนวนพิกเซล นี่คือสิ่งที่เครื่องได้รับเมื่อรถอยู่ไกล กล้องราคาถูก หรือภาพถูกบีบอัดมาก เมื่อพิกเซลน้อยลง สิ่งที่มันเห็นก็หายไปทีละอย่าง ทั้งที่รถยังจอดอยู่ตรงนั้น
Slide to take pixels away. This is what the machine gets when a car is far off, the camera is cheap, or the picture is squeezed hard. As pixels go, its findings vanish one by one — while the car is still standing right there.
- มันรู้เฉพาะสิ่งที่เคยถูกสอน COCO มี 80 ประเภท ภาพส่วนใหญ่มาจากเว็บ Flickr ไม่มีตุ๊กตุ๊ก ไม่มีสองแถว ไม่มีรถเข็นขายของ มันจะเรียกตุ๊กตุ๊กว่า “รถยนต์” หรือ “รถจักรยานยนต์” หรือไม่เห็นเลย ข้อมูลฝึกที่เอียง ทำให้เครื่องเอียงตาม
- กลางคืน ฝน หมอก แสงสะท้อน เลนส์สกปรก ทุกอย่างนี้เปลี่ยนตัวเลขทุกตัวในภาพ เมื่อสภาพต่างจากภาพที่มันเคยเห็นตอนฝึก มันจะพลาดมากขึ้น
- “ไม่เห็น” ไม่ได้แปลว่า “ไม่มี” อย่าถือว่าผลว่างเปล่าเป็นหลักฐานว่าถนนว่าง
- มั่นใจไม่ได้แปลว่าถูก 90% คือคะแนนที่มันให้ตัวเอง ไม่ใช่โอกาสถูก 90 ใน 100
- มันบอกได้ว่า “คน” แต่ไม่บอกว่าเป็นใคร เราไม่จดจำใบหน้าและไม่อ่านป้ายทะเบียน ไม่ว่ากล้องไหน งานวิจัยพบว่าระบบวิเคราะห์ใบหน้าเชิงพาณิชย์ผิดพลาดกับคนบางกลุ่มมากกว่ากลุ่มอื่นอย่างชัดเจน (Buolamwini และ Gebru, 2018) และภายใต้ พ.ร.บ.คุ้มครองข้อมูลส่วนบุคคล พ.ศ. 2562 ใบหน้าที่ระบุตัวคนได้คือข้อมูลส่วนบุคคล
- It only knows what it was taught. COCO has 80 kinds of thing, mostly from photos on Flickr. No tuk-tuk, no songthaew, no street-food cart. It will call a tuk-tuk a “car” or a “motorcycle”, or see nothing. Lopsided training data makes a lopsided machine.
- Night, rain, fog, glare, a dirty lens. Each changes every number in the picture. When conditions differ from the pictures it trained on, it misses more.
- “Not detected” does not mean “not there”. Never treat an empty result as proof that the road is empty.
- Confident is not correct. 90% is the score it gives itself, not a 90-in-100 chance of being right.
- It can say “person”, never who. We do not recognise faces or read number plates, on any camera. Research has shown commercial face-analysis systems making far more mistakes for some groups of people than others (Buolamwini & Gebru, 2018), and under Thailand's Personal Data Protection Act (2019) a face that identifies someone is personal data.
ต่อจากนี้Where next
คุณเห็นแล้วว่าเครื่องมองอย่างไร ต่อไปอ่านว่าทำไมมันถึงทำงาน สอนมันเอง หรือหาทางหลอกตามัน
You have seen how it looks. Now read why it works, teach one yourself, or find ways to fool it.
- HBคู่มือ: ทฤษฎีและแบบฝึกหัดHandbook: theory & exercisesอ่านทั้งแปดบทแบบเต็ม พร้อมแบบฝึกหัดให้ลองทำเอง คำถามท้ายบท และอภิธานศัพท์ไทย–อังกฤษAll eight chapters in full, with exercises you can try yourself, questions to check yourself, and a Thai–English glossary.
- 03ฝึกเครื่องด้วยตัวเองTrain a machine yourselfสอนด้วยตัวอย่างไม่กี่ภาพจากกล้องของคุณ ฝึกเสร็จในไม่กี่วินาทีTeach it with a few examples from your camera; training takes seconds.
- 04สั่งสิ่งเล็ก ๆ ด้วยท่ามือGestures that do thingsให้ท่ามือของคุณสั่งสิ่งเล็ก ๆ ในเบราว์เซอร์ พร้อมดูว่าเบราว์เซอร์ไม่ยอมให้ทำอะไรLet your hand run small things in the browser — and see what a browser will not allow.
- 05ขับเอง: รถที่บอกว่าเห็นอะไรDrive: cars that say what they seeสนามทดสอบมีวงเวียน ไฟจราจร ทางม้าลาย คน และสุนัขจร รถหกคันแข่งกันทำรอบโดยไม่ชนใคร และบอกเหตุผลทุกครั้งที่เบรกA test track with a roundabout, a light, zebras, people and stray dogs. Six cars race for laps without hitting anyone — and say why every time they brake.
- 06เกม: คนกับเครื่องGames: you vs the machineนับรถแข่งกับเครื่อง และทายภาพจากพิกเซลน้อยที่สุดOut-count it, and out-guess it from the fewest pixels.
- 10ข้อกฎหมายฉบับเต็มThe fine printทำไมเราไม่ระบุตัวบุคคล และกล้องเป็นของใครWhy we never identify anyone, and who owns the cameras.