VISION
ห้อง 03 · ฝึกเองRoom 03 · Train it yourself

ลองสอนเครื่องให้ทายภาพTeach it something new

วิดีโอเดียว เรียนรู้ได้สามแบบ เลือกวิธีด้านล่าง แล้วทำตามคำแนะนำข้างจอ เริ่มจากวิดีโอฝึกได้ทันที ไม่ต้องมีกล้องOne video. Three ways to learn. Choose a method below, then follow the next step beside the screen. Start with the practice video — no camera needed.

1 · เลือกวิดีโอ1 · Choose your video

วิดีโอฝึกสลับรถแน่นและถนนโล่งทุก 8 วินาที ถ้าใช้กล้องจริง ให้เลือกภาพที่เห็นความต่างชัด กล้องบางตัวส่งแค่ภาพนิ่ง เปลี่ยนได้เสมอThe practice clip alternates busy and quiet traffic every 8 seconds. For real cameras, choose a clear view with changing traffic. Some public cameras send still images; switch whenever you need.

2 · เลือกวิธีเรียนรู้2 · Choose how it learns

วิดีโอเดียวกันสำหรับทั้งสามวิธีThe same video for all three methods

กดชื่อที่ตรงกับภาพตอนนี้ อย่างน้อยกลุ่มละ 3 ภาพ แนะนำกลุ่มละ 10 ภาพ เว้นช่วงเพื่อให้ภาพต่างกันClick the name that matches this frame. Get at least 3 per group; aim for 10. Wait between clicks for different pictures.

ทุกวิธีใช้ MobileNet ที่ฝึกมาแล้วเพื่ออธิบายภาพ จากนั้นเรียนกฎเล็ก ๆ ของบทเรียนนี้ ไม่ได้ฝึกสมองใหญ่ใหม่ ไม่มีการตรวจนับรถหรือควบคุมรถจริง ข้อมูลอยู่ในแท็บนี้ ปิดหน้าแล้วหายAll three methods borrow a pretrained MobileNet to describe each frame, then learn a small rule for this lesson. They do not retrain the large network, count vehicles, or control real cars. Learning stays in this tab and disappears when you leave.

พร้อมลองต่อ? สอนสิ่งของของคุณ และดูข้างในโมเดลReady for more? Teach your own objects and look inside the model

อยากสอนอะไร?What would you like to teach?

05 → สอนท่ามือให้แล็ปท็อปTeach your laptop a gestureนิ้วเดียว ฝ่ามือ หรือไม่ทำท่า — แล้วให้แต่ละท่าทำอะไรสักอย่าง (ห้อง 04)One finger, a palm, or nothing — then make each one do something (room 04)

เริ่มที่ “ลองเลย” ได้ทันที หรือเลือกกิจกรรมที่ใช้กล้องของคุณStart with “Try it without a camera”, or choose a camera activity.

มันทำงานยังไงHow it works

Hand up 87% ยกมือ 87% A few examples ตัวอย่างไม่กี่ภาพ yours, from any camera ของคุณ จากกล้องใดก็ได้ A 1,280-number fingerprint ลายนิ้วมือ 1,280 ตัวเลข same picture, same numbers ภาพเดิม ได้เลขเดิมเสมอ Neighbours vote เพื่อนบ้านโหวต which examples are closest? ตัวอย่างไหนใกล้ที่สุด A guess + confidence คำเดา + ความมั่นใจ it always shows the number บอกคะแนนเสมอ
สี่ขั้นจากตัวอย่างของคุณสู่คำเดาที่บอกความมั่นใจเสมอFour steps from your examples to a guess that always shows its confidence.
01

ให้มันดูตัวอย่างShow it examples

ให้เครื่องดูภาพของแต่ละกลุ่ม กด “เพิ่มตัวอย่าง” ตอนที่เห็นสิ่งนั้น ลองเก็บกลุ่มละ 10 ภาพจากหลายมุม แล้วกด “ฝึกแล้วลองทาย”Show it each group. Press “Add example” when that thing is in the picture. Aim for 10 pictures per group from different angles. Then press “Train and try it”.

  1. เพิ่มตัวอย่างกลุ่มละอย่างน้อยหนึ่งภาพAdd examples to both groups
  2. ดูมันเริ่มเดาจากภาพสดWatch it start guessing from the live picture
  3. ฝึกแล้วลองภาพใหม่Train it, then try a new picture
  4. ถ้าอยากลองต่อ: ทดสอบกับกล้องถนนอื่นOptional: try other road cameras
ภาพที่เครื่องเห็นตอนนี้What it sees now
…

สองแถบนี้คือสองวิธีทายของเครื่อง แถบบนทำงานทันที ไม่ต้องฝึก แถบล่างมีชีวิตขึ้นมาก็ต่อเมื่อกด “ฝึกแล้วลองทาย” แล้วเท่านั้นThese two bars are two ways the machine guesses. The top one works at once, with no training. The bottom one comes alive only after you press “Train and try it”.

ทายทันที — ดูภาพที่คล้ายกันGuesses right away — from similar examples
ทายหลังกดฝึกGuesses after you press Train
ดูว่ามันทายผิดน้อยลงยังไงSee how its mistakes shrink

คะแนนบอกว่าภาพคล้ายกลุ่มไหน ไม่ได้แปลว่าเครื่องตอบถูกแน่นอน ลองของที่ไม่เคยสอน แล้วดูว่ามันสับสนไหมThe bars show which group looks most similar. A high score can still be wrong. Try something you never taught it and see what happens.

    ปุ่มนี้วาดภาพถนนจำลอง “รถแน่น / ถนนโล่ง” ขึ้นมาเอง 16 ภาพ — ไม่ใช่ภาพถ่าย ไม่ได้ใช้กล้อง และไม่มีอะไรออกจากเครื่องคุณ ใช้ดูให้ทุกชิ้นส่วนของหน้านี้ทำงานได้ในคลิกเดียวThis button draws 16 synthetic “busy / empty road” frames itself — not photographs, not from a camera, and nothing leaves your device. It exists so every instrument on this page can move in one click.

    อยู่ในแท็บนี้เท่านั้นThis tab only. ตัวอย่างทั้งหมดอยู่ในหน่วยความจำของแท็บนี้ ไม่ได้บันทึกและไม่ได้ส่งไปไหน โหลดหน้าใหม่เมื่อไรก็หายหมด ภาพจากกล้องของคุณประมวลผลในเครื่องนี้เท่านั้นYour examples live in this tab's memory — not saved, not sent anywhere. Reload the page and they are gone. Your camera's picture is processed on this device only.

    เริ่มที่กลุ่มละ 10–20 ภาพ และให้ตัวอย่างหลากหลาย ทั้งมุม ระยะ และแสง ถ้าทุกภาพของ “ยกมือ” คุณใส่เสื้อแดง เครื่องอาจเรียนรู้เสื้อแดงแทนมือStart with 10–20 per class and vary them: angle, distance, light. If you wear a red shirt in every “hand up” example, it may learn the shirt instead of the hand.

    ดูข้างใน: เครื่องเรียนรู้ยังไง (ไม่บังคับ)Look inside: how the machine learns (optional)

    การเรียนรู้แบบถ่ายโอนTransfer learning

    1. ภาพหนึ่งเฟรมจากกล้องOne frame from the camera
    2. MobileNetV2 — แช่แข็งไว้ ไม่ได้ฝึกใหม่MobileNetV2 — frozen, never retrained
    3. ตัวเลข 1,280 ตัว ลายนิ้วมือของภาพ1,280 numbers: the picture's fingerprint
    4. การตัดสินใจของคุณ — ส่วนเดียวที่เรียนรู้Your decision — the only part that learns

    โครงข่ายนี้ถูกฝึกมาแล้วกับภาพจาก ImageNet กว่าล้านภาพ (Sandler และคณะ, 2018) ระหว่างทาง ชั้นต่าง ๆ ของมันเรียนรู้ที่จะเห็นขอบ ลวดลาย พื้นผิว และส่วนประกอบของสิ่งของ เราขอยืมความรู้นั้นมาทั้งก้อน แล้วฝึกเพียงชั้นสุดท้ายเล็ก ๆ ที่มีน้ำหนัก (ตัวเลขในโมเดลที่การฝึกคอยปรับ) แค่กลุ่มละ 1,280 ตัว บวกอีกหนึ่งตัว ตัวอย่างไม่กี่สิบภาพจึงพอ และฝึกเสร็จในไม่กี่วินาทีแม้บนโทรศัพท์

    The network was already trained on more than a million ImageNet photos (Sandler et al., 2018). Along the way its layers learned to pick out edges, textures, patterns and parts of things. We borrow all of that as it is and train only a small last layer: 1,280 weights per class, plus one. A weight is one of the adjustable numbers inside the model that training nudges. That is why a few dozen examples are enough, and why training takes seconds, even on a phone.

    02

    สองวิธีเรียนรู้Two ways to learn

    เพื่อนบ้านใกล้สุด — ไม่ต้องฝึกเลยNearest neighbours — no training at all

    จำตัวอย่างไว้ทุกภาพ เมื่อเจอภาพใหม่ ก็หาตัวอย่าง 5 ภาพที่ตัวเลข 1,280 ตัว “ชี้ไปทางเดียวกัน” มากที่สุด แล้วให้มันโหวต ภาพที่คล้ายกว่าได้เสียงมากกว่า วิธีนี้เก่าแก่มาก (Cover และ Hart, 1967) และใช้ได้ทันทีที่คุณเพิ่มตัวอย่าง

    Remember every example. For a new picture, find the 5 examples whose 1,280 numbers point most nearly the same way, and let them vote — closer ones count more. The idea is old (Cover & Hart, 1967) and works the instant you add an example.

    ทดสอบแบบเว้นทีละภาพLeave-one-out test

    —

    ซ่อนตัวอย่างไว้หนึ่งภาพ ให้ภาพที่เหลือทายว่าเป็นกลุ่มไหน ทำแบบนี้ทีละภาพจนครบ สัดส่วนที่ทายถูกคือการเดาที่ซื่อตรงกว่าว่ามันจะทำได้ดีแค่ไหนกับภาพที่ไม่เคยเห็น แต่ก็ยังเป็นภาพแบบเดียวกับที่คุณถ่ายเท่านั้นHide one example, let the rest label it, and repeat for every example. The share it gets right is a fairer guess at how it will do on pictures it has not seen — though still only pictures like the ones you took.

    ชั้นที่ฝึกขึ้นมา — ปรับทีละก้าวA trained layer — gradient descent

    ชั้นน้ำหนักเดียว (softmax regression — สูตรที่เปลี่ยนตัวเลขดิบเป็นคะแนนของแต่ละกลุ่ม) เริ่มจากค่าสุ่ม แล้วถูกปรับทีละนิด ซ้ำ ๆ ไปเรื่อย ๆ แบบที่เรียกว่าการลดตามความชัน (gradient descent) คือปรับตัวเลขทุกตัวเล็กน้อยในทิศทางที่ทำให้เดาผิดน้อยลง โดยมีวิธี Adam (Kingma และ Ba, 2015) คอยกำหนดว่าแต่ละครั้งจะปรับมากน้อยแค่ไหน จนแยกกลุ่มของคุณออกจากกันได้ กด “ฝึก” แล้วดูเส้นค่าผิดพลาดค่อย ๆ ลดลง ตลอด 200 รอบที่มันไล่ดูตัวอย่างทั้งหมด

    One layer of weights (softmax regression — a formula that turns raw numbers into a score for each group) starts random and is nudged a little at a time, pass after pass, in the way called gradient descent: adjust every number a little in whichever direction makes it wrong less often. Adam (Kingma & Ba, 2015) decides how big each nudge is, until it separates your classes. Press Train and watch the loss — one number saying how wrong it is right now — fall, over 200 passes of every example.

    ใช้ปุ่มฝึกใต้จอ แล้วลองภาพใหม่เพื่อดูว่ามันเรียนรู้อะไรUse the training button below the demonstration. Then try a new picture to see what it learned.

    ความแม่นตรงนี้วัดจากตัวอย่างชุดเดียวกับที่มันใช้เรียน จึงดูดีเกินจริง คะแนนแบบเว้นทีละภาพทางซ้ายเป็นการทดสอบที่ยุติธรรมกว่าAccuracy here is measured on the very examples it learned from, so it flatters. The leave-one-out score is the fairer test.

    ตัวเลข 1,280 ตัวของตัวอย่างแต่ละภาพถูกบีบเหลือ 2 ตัวด้วยการวิเคราะห์องค์ประกอบหลัก (PCA) รูปทรงบอกกลุ่ม วงสีส้มคือภาพตอนนี้ เส้นบางชี้ไปยังเพื่อนบ้านที่กำลังโหวต นี่เป็นเพียงเงา จุดที่ใกล้กันใน 1,280 มิติอาจดูห่างกันบนแผนที่นี้Each example's 1,280 numbers squeezed to 2 by principal component analysis (PCA). Shapes mark classes; the orange ring is the picture now; thin lines run to the neighbours currently voting. It is only a shadow — points close in 1,280 dimensions can look far apart here.
    03

    ลองกับกล้องถนนอื่นTry it on other road cameras

    สอนเรื่องถนนแล้วใช่ไหม? ลองกล้องถนนอื่นอีก 24 ตัว แล้วดูว่ามันสับสนตรงไหน ถ้าสอนเรื่องมือหรือแก้ว ให้ลองภาพใหม่ของมือหรือแก้วก่อน ภาพถนนไม่ใช่ข้อสอบที่เหมาะกับสิ่งที่สอน

    Taught it about roads? Try 24 other road cameras. Watch where it gets confused. If you taught it about hands or cups, road pictures will not be a fair test — try new pictures of those things first.

    ป้ายสีส้มคือคำตอบที่มั่นใจตั้งแต่ 70% ขึ้นไป ต่ำกว่านั้นเราเขียนว่า “ไม่แน่ใจ”An orange tag means 70% or more for one class; below that we write “unsure”.

    04

    เก็บไว้ หรือปล่อยให้ลืมKeep it, or let it forget

    อยากใช้สิ่งที่สอนอีกครั้งไหม? บันทึกเป็นไฟล์ไว้ แล้วเปิดกลับมาในหน้านี้ได้ ไฟล์เก็บชื่อกลุ่มและสิ่งที่เครื่องเรียนรู้ แต่ไม่เก็บรูปของคุณ และไม่มีอะไรถูกส่งไปที่อื่น

    Want to use what you taught next time? Save it as a file, then open it here later. The file keeps the group names and what the computer learned. It does not keep your pictures. Nothing is uploaded.