VISION
ห้อง 08 · งานวิจัยและเรื่องของเว็บนี้Room 08 · Research and about

หกสิบปี
ของการสอนเครื่อง
ให้มองเห็น
Sixty years
of teaching
machines to see

จากภาพถ่ายบล็อกไม้ในปี 1963 ถึงแบบจำลองที่ตัดขอบอะไรก็ได้ที่คุณชี้ นี่คืองานวิจัยที่เปลี่ยนวิธีคิด งานที่เว็บนี้ใช้อยู่จริง และคำถามที่ยังไม่มีใครตอบได้

From a photograph of blocks in 1963 to a model that outlines anything you point at: the papers that changed the thinking, the ones this site actually runs, and the questions nobody has answered yet.

ทุกงานในหน้านี้มีอยู่จริง เราอ้างผู้เขียน ปี และที่ตีพิมพ์ ถ้าไม่แน่ใจในรายละเอียดใด เราตัดทิ้ง ชื่องานคงไว้เป็นภาษาอังกฤษตามต้นฉบับEvery work on this page is real. We give authors, year and venue; where we were not sure of a detail, we left it out.

—กล้องสาธารณะในรายชื่อpublic cameras listed
—เครื่องอ่านพิกเซลได้a machine can read
—วิดีโอสดlive video
—ภาพนิ่งstills
—หน่วยงานเจ้าของกล้องagencies that own them
—ออฟไลน์ตอนตรวจล่าสุดoffline at last check
00

ห้องเรียนนี้คืออะไร และยืนอยู่บนหลักฐานอะไรWhat this classroom is, and what it stands on

ตัวเลขข้างบนไม่ใช่ตัวเลขตกแต่ง มันคือรายชื่อกล้องที่เว็บนี้อ่านอยู่จริงตอนนี้The numbers above are not decoration. They are the camera list this site is reading right now.

เว็บนี้คือห้องเรียนสาธารณะเรื่องคอมพิวเตอร์วิทัศน์ ยืนอยู่บนกล้องสาธารณะของประเทศไทยเอง: กล้องหลายพันตัวของหน่วยงานต่าง ๆ ที่เปิดให้ประชาชนดูได้ เรารวมรายชื่อจากFloodDash และพยายามอัปเดตทุกสิบนาที โมเดลทุกตัวรันในเบราว์เซอร์ของคุณ ไม่มีภาพถูกเก็บ และไม่มีใครถูกระบุตัว — นี่คือเส้นที่เราวาดไว้ตั้งแต่วันแรก กล้องทุกตัวเป็นของหน่วยงานเจ้าของที่ติดตั้งและดูแล เราแค่ลิงก์ไปและส่งต่อบางภาพ

This site is a public classroom for computer vision, standing on Thailand's own public cameras: thousands of them, run by the agencies that installed them and opened to the public. We gather the list through FloodDash and try to refresh it every ten minutes. Every model runs in your browser, no picture is stored, and nobody is identified — a line we drew on day one. Every camera belongs to the agency that installed and runs it; we only link to them and pass some frames along.

สร้างและดูแลโดย ดร.นน อัครประเสริฐกุล — เรื่องเบื้องหลังว่าทำไมมันเป็นแบบนี้ · หน้าระบบว่าอะไรรันอยู่ใต้ฝา · หน้ากล้องที่ทำให้ตัวเลขข้างบนมีชีวิต · และ กติกาที่เราถือ หน้าที่คุณกำลังอ่านคือหน้าวิจัยของเว็บ: เริ่มจากสี่ก้าวสำคัญของประวัติศาสตร์ แล้วไล่ต่อไปถึงงานที่เว็บนี้ใช้จริง และคำถามที่ยังไม่มีใครตอบได้

Made and run by Dr Non Arkaraprasertkul — the story of why it looks like this · the system page for what runs under the bonnet · the cameras page, where the numbers above come alive · and the rules we hold. The page you are reading is the site's research-and-about room: it starts at four landmarks of the field, then walks down to the work this site actually runs and the questions nobody has answered.

จากเส้น สู่คำทาย: ลองเดินผ่านประวัติศาสตร์From lines to guesses: walk through the history

กดปีเพื่อดูว่าคำถามและวิธีคิดเปลี่ยนไปอย่างไร นี่คือสี่ตัวอย่างสำคัญ ไม่ใช่ทุกก้าวของประวัติศาสตร์ วิธีเก่ากับวิธีใหม่ยังใช้ร่วมกันได้Choose a year to see how the question and approach changed. These are four landmarks, not the whole history. Older and newer methods still work together.

ภาพอธิบายที่วาดใหม่ ไม่ใช่ผลจากโมเดลOriginal explanatory illustration, not model output
1963

เส้นในภาพบอกอะไร?What do the lines tell us?

ในภาพนี้ สีแต่ละด้านช่วยให้เราเห็นลูกบาศก์ นักวิจัยยุคแรกให้เครื่องหาขอบ แล้วใช้กฎที่คนเขียนจับคู่กับรูปทรงที่รู้จัก งานของ Roberts ศึกษาฉากที่มีทรงเรขาคณิตง่าย ๆThe coloured faces help us see a cube. Early researchers used edges and rules to match familiar solid shapes. Roberts studied scenes of simple geometric objects.

ฉากเรียบง่ายช่วยให้ทำได้ แต่ต้นไม้ เงา หรือของที่บังกันทำให้กฎยุ่งยากขึ้นSimple scenes make this easier. Trees, shadows and overlapping objects make rules much harder.

Roberts · MIT thesis · 1963 · ดาวน์โหลดภาพอธิบายDownload illustration

ภาพอธิบายที่วาดใหม่ ไม่ใช่ผลจากโมเดลOriginal explanatory illustration, not model output
2005

รูปร่างนี้คล้ายคนไหม?Does this shape look like a person?

รูปคนทางซ้ายถูกย่อเป็นทิศทางของขอบและจำนวนทางขวา HOG ของ Dalal และ Triggs วัดว่าขอบในแต่ละบริเวณชี้ไปทางไหน แล้วใช้ตัวจำแนกที่ฝึกไว้ตัดสินว่าเป็นคนหรือไม่The person on the left becomes edge directions and counts on the right. Dalal and Triggs’s HOG measured local edge directions. A trained classifier used those measurements to decide whether a window contained a person.

คนยังเลือกว่าจะวัดอะไร เครื่องอาจพลาดเมื่อภาพต่างจากตัวอย่างที่ใช้ฝึกPeople still choose what to measure. Different poses or pictures can confuse a system trained on other examples.

Dalal & Triggs · CVPR · 2005 · ดาวน์โหลดภาพอธิบายDownload illustration

ภาพอธิบายที่วาดใหม่ ไม่ใช่ผลจากโมเดลOriginal explanatory illustration, not model output
2016

มีอะไรอยู่ตรงไหน?What is here, and where?

กรอบล้อมบริเวณที่คาดว่ามีรถ โครงข่ายเรียนรู้ลักษณะที่มีประโยชน์จากภาพฝึก แทนให้คนออกแบบทุกลักษณะ YOLO รวมการทายชนิดและกรอบไว้ในโครงข่ายเดียวThe box marks where a car is thought to be. A neural network learns useful features from training pictures. YOLO combined guesses about object types and their boxes in one network.

กรอบเป็นคำทาย ไม่ใช่หลักฐานว่าเครื่องเข้าใจรถ ภาพเบลอหรือวัตถุเล็กอาจทำให้พลาด เว็บนี้ใช้ SSDLite ไม่ใช่ YOLOA box is a guess, not proof that it understands cars. Blur and small objects can cause misses. This site uses SSDLite, not YOLO.

Redmon et al. · CVPR · 2016 · ดาวน์โหลดภาพอธิบายDownload illustration

ภาพอธิบายที่วาดใหม่ ไม่ใช่ผลจากโมเดลOriginal explanatory illustration, not model output
2023

พิกเซลไหนเป็นส่วนของสิ่งนี้?Which pixels belong to this object?

สีส้มตามรูปรถละเอียดกว่ากรอบ นี่เรียกว่าการแบ่งส่วนภาพ SAM รับจุดหรือกรอบเป็นคำใบ้ แล้วสร้างพื้นที่ที่คาดว่าเป็นวัตถุที่ต้องการOrange follows the car’s shape more closely than a box. This is segmentation. SAM takes prompts such as points or boxes and proposes a mask for the selected object.

พื้นที่ที่เลือกอาจยังผิด และไม่ได้บอกโดยอัตโนมัติว่าวัตถุนั้นคืออะไร เว็บนี้ไม่ได้รัน SAM ภาพนี้เป็นเพียงภาพอธิบายThe selected area can still be wrong. A mask does not automatically name the object. This site does not run SAM; this is an explanatory drawing.

Kirillov et al. · ICCV · 2023 · ดาวน์โหลดภาพอธิบายDownload illustration

เครื่องเรียนรู้ได้สามแบบThree ways a machine can learn

คำถามสำคัญคือ เครื่องได้รับคำตอบแบบไหน? บางครั้งเราให้ชื่อที่ถูกต้อง บางครั้งให้แค่ข้อมูล และบางครั้งให้คะแนนหลังจากมันลองทำThe key question is: what feedback does the machine get? Sometimes we give it correct answers, sometimes just data, and sometimes a reward after it tries an action.

01

มีคำตอบให้: Supervised learningLearn with answers: supervised learning

เหมือนฝึกจากบัตรคำที่มีเฉลย เราให้ภาพพร้อมชื่อ เช่น “วงกลม” หรือ “สี่เหลี่ยม” เครื่องใช้ตัวอย่างเพื่อเรียนรู้การทายชื่อของภาพใหม่Like practising with flashcards that have answers. We give each picture a label, such as “circle” or “square”. The machine uses these examples to learn how to label a new picture.

  • วงกลมสีส้มOrange circle
  • สี่เหลี่ยมสีเหลืองYellow square
  • วงกลมสีเหลืองYellow circle
  • สี่เหลี่ยมสีส้มOrange square

ต้องลองภาพที่ไม่ได้ใช้ฝึกด้วย จำเฉลยชุดเดิมได้ ไม่ได้แปลว่าทายภาพใหม่ได้เสมอTest with pictures it did not learn from. Remembering the practice answers does not guarantee success on new pictures.

ลองสอนด้วยตัวอย่างในห้องฝึก →Teach with examples in Training →
02

ไม่มีเฉลย: Unsupervised learningFind patterns: unsupervised learning

เราให้ข้อมูลโดยไม่บอกชื่อที่ถูกต้อง เครื่องมองหาความคล้ายกัน เช่น จัดภาพเป็นกลุ่ม สีหรือรูปร่างอาจทำให้ได้กลุ่มคนละแบบ กลุ่มเหล่านี้ไม่ได้มีชื่อว่า “แมว” หรือ “สุนัข” ขึ้นมาเองWe give data without the correct labels. The machine looks for patterns, such as groups of similar pictures. Colour and shape can produce different groups. These groups do not automatically name themselves “cat” or “dog”.

  • วงกลมสีส้มOrange circle
  • สี่เหลี่ยมสีเหลืองYellow square
  • วงกลมสีเหลืองYellow circle
  • สี่เหลี่ยมสีส้มOrange square

ภาพอธิบาย: เราเลือกกฎจัดกลุ่มง่าย ๆ เพื่อให้เห็นผลของการเลือกความคล้าย นี่ไม่ใช่การฝึกโมเดลจัดกลุ่มGuided illustration: these buttons use simple sorting rules to show how a choice of similarity changes the groups. They do not train a clustering model.

พบรูปแบบได้ แต่ต้องตรวจว่ารูปแบบนั้นมีประโยชน์หรือเป็นเพียงสิ่งบังเอิญFinding a pattern does not tell us whether it is useful or just a coincidence.

03

ลองทำแล้วรับคะแนน: Reinforcement learningTry, get a reward: reinforcement learning

เหมือนหุ่นยนต์ลองเลือกทาง มันทำบางอย่าง รับคะแนน แล้วปรับว่าครั้งต่อไปควรเลือกอะไร เป้าหมายคือได้คะแนนรวมดีขึ้น ไม่ใช่รับเฉลยให้ทุกการกระทำLike a robot trying a route. It takes an action, gets a reward, then adjusts what to choose next. Its goal is to earn better rewards, rather than receive the correct answer for every action.

← ทางซ้าย: ไปถึงเป้าหมายLeft path: reach the goal+1
→ ทางขวา: เจอสิ่งกีดขวางRight path: hit an obstacle−1

ตัวอย่างเรียนรู้จริงแบบหนึ่งก้าว: จำคะแนนเฉลี่ยของสองทาง เริ่มจากลองทางที่ยังไม่รู้ แล้วเลือกทางที่เคยได้คะแนนดีกว่า ระบบที่ต้องเดินหลายก้าวซับซ้อนกว่านี้มากA real, tiny one-step reward learner: it remembers each path’s average reward, tries unknown paths first, then chooses the better one. Learning sequences of many actions is much more complex.

คะแนนต้องตรงกับสิ่งที่เราต้องการ ถ้าให้คะแนนความเร็วอย่างเดียว หุ่นยนต์อาจเรียนรู้ให้รีบโดยไม่ปลอดภัย สนามขับรถในเว็บนี้ใช้กฎที่เขียนไว้ ไม่ได้ฝึกด้วย reinforcement learningRewards must match what we want. Rewarding only speed can encourage unsafe behaviour. This site’s driving track uses written rules; it is not trained with reinforcement learning.

จำง่าย ๆ: มีเฉลย → เรียนจากตัวอย่าง; ไม่มีเฉลย → หารูปแบบ; มีคะแนนหลังทำ → เรียนจากผลของการกระทำ คำว่า “non-supervised” มักหมายถึง “unsupervised” วิธีเหล่านี้ใช้ร่วมกันได้ และโมเดลที่เรียนรู้เองจากคำตอบที่สร้างจากข้อมูลเรียกว่า self-supervised ซึ่งเป็นอีกแนวคิดหนึ่งRemember: answers → learn from examples; no answers → find patterns; rewards after actions → learn from outcomes. “Non-supervised” usually means “unsupervised”. These approaches can work together. Self-supervised learning, where training targets come from the data itself, is a related but distinct idea.

กรณีศึกษา · ประเทศไทยCASE STUDY · THAILAND

NT TAG ID

จากภาพกล้อง สู่การจับคู่อัตลักษณ์From camera pixels to identity matching

เว็บไซต์โครงการ ↗Project website ↗

NT TAG ID อธิบายระบบที่รวมการวัดลักษณะภาพแบบดั้งเดิมกับการเรียนรู้เชิงลึก (deep learning: เรียนรู้จากภาพจำนวนมาก แทนที่จะเขียนกฎเอง) แปลงภาพเป็นรูปแบบตัวเลข แล้วใช้จับคู่บุคคลข้ามกล้อง ระบบนี้ทำให้เห็นว่าการตรวจพบใบหน้า การติดตาม และการระบุตัวบุคคล เป็นคนละโจทย์กัน

NT TAG ID describes a system combining classical image features with deep learning (learning from many examples instead of rules people write), encoding images as numerical patterns for matching people across cameras. It brings three different research tasks into view: detecting a face, following a track, and matching an identity.

อ่านสถาปัตยกรรมแบบคู่Reading a dual architecture
01ภาพจากกล้องCamera frameตรวจหาและจัดแนวบริเวณใบหน้าDetect and align the face region
เส้นทางดั้งเดิมClassical stream วัดสิ่งที่คนออกแบบMeasure designed features

รูปร่าง ขอบ และพื้นผิวของภาพShape, edges, and image texture

เส้นทางเรียนรู้เชิงลึกDeep stream เรียนรู้คำอธิบายจากพิกเซลLearn a representation from pixels

โครงข่ายสร้างเวกเตอร์ลักษณะภาพ (ชุดตัวเลขที่ใช้แทนภาพหนึ่งภาพ)A network produces a feature vector (a list of numbers standing for one picture)

02 → 03รวมหลักฐาน → ตัดสินการจับคู่Combine evidence → decide a matchการแปลงเป็นตัวเลขยังไม่ใช่คำตอบ ต้องมีวิธีเปรียบเทียบและเกณฑ์ตัดสินEncoding is one step; comparison and a decision threshold still matter

ภาพอธิบายที่วาดขึ้นใหม่จากคำอธิบายสถาปัตยกรรมคู่บนเว็บไซต์ NT TAG ID

An explanatory drawing based on NT TAG ID’s published dual-architecture description.

เจาะงานวิจัยRESEARCH LENS

HFFL · การผสานแบบปรับตัวAdaptive fusion

แผนภาพวิทยานิพนธ์ที่ระบุชื่อ D. Vongkomolshet (2026) วาง HFFL ไว้ในกลุ่มวิธีผสม มีเส้นทางดั้งเดิม เส้นทางโครงข่าย และส่วนผสานที่ปรับตามความมั่นใจและทรัพยากรประมวลผล เปิดแต่ละส่วนเพื่อดูว่าต่อยอดงานใด

The dissertation diagrams credited to D. Vongkomolshet (2026) position HFFL as a hybrid: a classical stream, a neural stream, and fusion that responds to confidence and computing resources. Open each part to see its research lineage.

01ลักษณะที่ออกแบบDesigned features

แผนภาพอ้าง Haar cascade ของ Viola–Jones สำหรับการตรวจหาใบหน้า และ HOG กับ LBP สำหรับคำอธิบายรูปร่างและพื้นผิว การหาใบหน้าเป็นขั้นแรก ไม่ได้บอกว่าเป็นใคร

The diagrams cite a Viola–Jones Haar cascade for face detection, with HOG and LBP describing shape and texture. Finding a face is a first step; it does not establish identity.

02ลักษณะที่เรียนรู้Learned features

เส้นทางโครงข่ายอ้าง YOLO สำหรับตรวจหา และ ArcFace สำหรับเวกเตอร์ใบหน้า (ชุดตัวเลขที่ใช้แทนใบหน้าหนึ่งใบ) โดยมีการจัดแนว ความสนใจ (กลไกที่ชั่งน้ำหนักว่าส่วนไหนควรดูก่อน) และการติดตามต่อเนื่องข้ามเฟรมประกอบอยู่ด้วย

The neural path cites YOLO detection and ArcFace face embeddings (a list of numbers standing for one face), alongside alignment, attention (a way of weighting which part to look at first), and continuity across video frames.

03ส่วนผสาน AFFOAFFO fusion

แผนภาพเสนอการถ่วงน้ำหนักแบบเปลี่ยนได้ โดยใช้ความมั่นใจ ความไม่แน่นอนแบบ entropy (ยิ่งคำตอบกระจาย ยิ่งไม่แน่ใจ) และความพร้อมของ GPU (การ์ดกราฟิกที่ใช้ประมวลผล) แทนการให้น้ำหนักสองเส้นทางเท่ากันตลอด

The diagrams propose changing the streams’ weights using confidence, entropy-based uncertainty (how spread out the scores are), and GPU availability (whether a graphics chip is free), rather than keeping a fixed mixture.

04ส่วนเสริม HFFL-TMHFFL-TM extension

เส้นทางเสริมใช้จุดสังเกตใบหน้า 68 จุด สร้างเมทริกซ์ระยะห่างและรหัสเลขฐานสอง การแยกเป็นสามเหลี่ยมถูกระบุว่าเป็นข้อกำหนดสำหรับงานต่อไป ยังไม่ได้ทำในโค้ด

An optional path uses 68 facial landmarks, pairwise distances, and a binary code. Triangle decomposition is explicitly marked as a future specification, not yet implemented.

ที่มา: แผนภาพ “Five Eras of Face Recognition Theory” และ “Direct Contribution Model” ที่ระบุชื่อ D. Vongkomolshet เอกสารเหล่านี้อธิบายงานวิทยานิพนธ์ ไม่ได้ยืนยันว่า HFFL ทุกส่วนเป็นระบบ NT TAG ID ที่เปิดใช้งานอยู่ หรือผ่านการประเมินอิสระแล้ว

Source: “Five Eras of Face Recognition Theory” and “Direct Contribution Model” diagrams credited to D. Vongkomolshet. They describe dissertation work; they do not establish that every HFFL component is deployed in NT TAG ID or independently evaluated.

จากเรขาคณิต สู่การผสานแบบปรับตัวFrom geometry to adaptive fusion
  1. ยุคแรกEarly familiesเรขาคณิตและภาพรวมGeometry & whole imagesวัดระยะ หรืออ่านทั้งภาพเป็นรูปแบบMeasure distances or represent the whole image
  2. 1996–20051996–2005ลักษณะที่ออกแบบDesigned descriptorsพื้นผิว LBP · รูปร่าง HOG · การตรวจหาแบบ HaarLBP texture · HOG shape · Haar detection
  3. 2014–20152014–2015เวกเตอร์ที่เรียนรู้Learned embeddingsDeepFace และ FaceNet: เรียนรู้คำอธิบายใบหน้าDeepFace and FaceNet: learn a face representation
  4. 2017–20192017–2019ขอบเขตเชิงมุมAngular marginsSphereFace · CosFace · ArcFace: แยกเวกเตอร์ให้ชัดขึ้นSphereFace · CosFace · ArcFace: separate embeddings
  5. งานที่เสนอ · 2026Proposed work · 2026การผสานแบบปรับตัวAdaptive fusionHFFL: รวมวิธีดั้งเดิมกับวิธีเรียนรู้HFFL: combine designed and learned features

แผนผังแนวคิดที่วาดขึ้นใหม่จากแผนภาพวิทยานิพนธ์ วิธีเหล่านี้ยังใช้อยู่ร่วมกัน ไม่ใช่ยุคที่แทนกันทั้งหมด อ่านงานต้นฉบับ: DeepFace (2014) · FaceNet (2015) · ArcFace (2019)

A conceptual map redrawn from the dissertation diagrams. These families coexist; each did not replace the previous one. Read the original papers: DeepFace (2014) · FaceNet (2015) · ArcFace (2019).

ดูระบบจากผู้สร้างWatch the system’s own demos

วิดีโอจากช่อง @nttagid ของโครงการ ภาพบนปุ่มเป็นภาพอธิบายที่วาดขึ้นใหม่ ไม่ใช่ภาพจากวิดีโอ

Videos from the project’s @nttagid channel. The link artwork is an original illustration, not footage.

อ่านหลักฐานให้เป็นHow to read the evidence

คำอธิบายผลิตภัณฑ์บอกว่าระบบตั้งใจทำอะไร ส่วนการประเมินต้องบอกด้วยว่าทดสอบกับใคร ภายใต้แสงและมุมแบบไหน ผิดพลาดในการจับคู่และปฏิเสธเท่าไร และเวลาที่รายงานรวมการตรวจหาและการค้นหาหรือไม่ เราจึงนำเสนอเป็นกรณีศึกษา และไม่ยกคำกล่าวอ้างเรื่องความแม่นยำหรือความเร็วมาใช้แทนผลทดสอบจากภายนอก

A product description explains an intended system. Evaluation also needs the details: who was tested, under what light and at what angle, how often the system matched the wrong person or missed the right one, and which steps a reported time actually covers. We therefore present this as a case study, and we do not repeat an advertised accuracy or speed as though it were an independent test result.

การแทนใบหน้าด้วยตัวเลขไม่ได้ทำให้ความเชื่อมโยงกับบุคคลหายไปเอง หากตัวเลขยังใช้จับคู่คนเดิมได้ ก็ต้องศึกษาความเป็นส่วนตัวของตัวแทนนั้นด้วย ห้องเรียนนี้อธิบายงานระบุตัวบุคคล แต่เครื่องมือของเรายังคงนับสิ่งของโดยไม่ระบุตัวคน

Encoding a face as numbers does not by itself remove its connection to a person. A representation that matches the same person still needs privacy analysis. This classroom explains identity research; its instruments continue counting objects without identifying people.

01 · 1959–1998

กฎที่คนเขียนเองRules written by hand

คนบอกเครื่องว่าขอบคืออะไร รูปร่างคืออะไร แล้วหวังว่าโลกจะเรียบง่ายพอPeople told the machine what an edge is and what a shape is, and hoped the world would be simple enough.

  1. 1959

    Receptive fields of single neurones in the cat's striate cortex

    David H. Hubel, Torsten N. Wiesel · The Journal of Physiology

    เซลล์ประสาทในสมองส่วนการมองเห็นของแมวตอบสนองต่อขอบที่เอียงบางมุม ความคิดที่ว่าตัวตรวจจับง่าย ๆ ซ้อนกันเป็นตัวตรวจจับที่ซับซ้อนขึ้น กลายเป็นภาพตั้งต้นของคอมพิวเตอร์วิทัศน์ และต่อมาคือโครงข่ายคอนโวลูชัน

    Neurons in a cat's visual cortex fire for edges at particular angles. Simple detectors stacked into more complex ones became the founding picture of machine vision — and, later, of convolutional networks.

  2. 1963

    Machine Perception of Three-Dimensional Solids

    Lawrence G. Roberts · PhD thesis, MIT

    “โลกของบล็อก” จากภาพถ่ายทรงเรขาคณิตง่าย ๆ หาขอบ จับคู่กับรูปทรงสามมิติที่รู้จัก แล้วสร้างฉากกลับขึ้นมา ได้ผลเพราะโลกในภาพมีแต่บล็อก

    The blocks world. From a photograph of simple solids, find the edges, match them to known 3D shapes, rebuild the scene. It worked because the world in the picture was made of blocks.

  3. 1966

    The Summer Vision Project

    Seymour Papert · MIT AI Memo 100

    แผนให้นักศึกษาช่วงปิดภาคฤดูร้อนสร้างส่วนสำคัญของระบบการมองเห็นให้เสร็จ คนจำงานนี้ได้เพราะมันประเมินความยากต่ำไปมาก ปัญหานี้ยังไม่จบจนทุกวันนี้

    A plan for summer students to build a significant part of a visual system. It is remembered for how badly it underestimated the problem, which is still open.

  4. 1979

    A Threshold Selection Method from Gray-Level Histograms

    Nobuyuki Otsu · IEEE Transactions on Systems, Man, and Cybernetics ใช้ในเว็บนี้Used here

    เลือกค่าขีดแบ่งที่แยกพิกเซลมืดกับสว่างได้ขาดที่สุด ยังเป็นปุ่ม “อัตโนมัติ” ในเครื่องมือนับไม่ถ้วน รวมถึงเครื่องมือของเรา

    Choose the threshold that best separates dark pixels from light. Still the “auto” button in countless tools, including ours.

  5. 1980

    Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position

    Kunihiko Fukushima · Biological Cybernetics

    ชั้นของเซลล์แบบง่ายและแบบซับซ้อนตามแนวคิดของ Hubel กับ Wiesel ที่จำรูปแบบได้ไม่ว่าจะอยู่ตรงไหนของภาพ คือโครงสร้างของโครงข่ายคอนโวลูชัน ก่อนจะมีวิธีฝึกที่ดี

    Layers of simple and complex cells, after Hubel and Wiesel, that recognise a pattern wherever it appears. The shape of a convolutional network, before there was a good way to train one.

  6. 1982

    Vision: A Computational Investigation into the Human Representation and Processing of Visual Information

    David Marr · W. H. Freeman (book, published after his death in 1980)

    มองการมองเห็นเป็นสามระดับ: คำนวณอะไรและเพื่ออะไร คำนวณอย่างไร และบนฮาร์ดแวร์แบบไหน การเห็นคือการสร้างคำอธิบาย จากขอบ สู่ภาพร่าง สู่รูปทรงสามมิติ เป็นกรอบคิดที่นักวิจัยทั้งรุ่นใช้ถกเถียง

    Vision at three levels: what is computed and why, how, and in what hardware. Seeing as building descriptions — from edges, to a sketch, to 3D shape. The framework a generation of researchers argued with.

  7. 1986

    A Computational Approach to Edge Detection

    John Canny · IEEE Transactions on Pattern Analysis and Machine Intelligence

    ตัวหาขอบที่ออกแบบจากเป้าหมายที่ประกาศไว้ชัด: เจอขอบจริง วางตำแหน่งแม่น และทำเครื่องหมายขอบละครั้งเดียว เลนส์ “ขอบ” ของเราใช้ตัวดำเนินการ Sobel ที่เก่ากว่าและง่ายกว่า ส่วนของ Canny คือสิ่งที่ใช้เมื่อความแม่นสำคัญ

    Edge detection derived from stated goals: find real edges, place them accurately, mark each one once. Our Edges lens uses the older, simpler Sobel operator; Canny's is what you reach for when it matters.

  8. 1989

    Backpropagation Applied to Handwritten Zip Code Recognition

    Yann LeCun et al. · Neural Computation

    โครงข่ายคอนโวลูชันที่ฝึกทั้งระบบด้วย backpropagation (ย้อนกลับจากข้อผิดพลาดเพื่อปรับตัวเลขทีละนิด) อ่านรหัสไปรษณีย์ลายมือบนจดหมาย ต่อมา LeNet-5 (LeCun, Bottou, Bengio, Haffner, Proceedings of the IEEE, 1998) ถูกนำไปใช้อ่านเช็คจริง

    A convolutional network trained end to end with backpropagation (working backwards from the mistake to adjust each number a little), reading handwritten zip codes on mail. Its successor, LeNet-5 (LeCun, Bottou, Bengio, Haffner, Proceedings of the IEEE, 1998), went on to read real cheques.

02 · 1999–2011

ลักษณะที่คนออกแบบFeatures designed by hand

คนออกแบบว่าจะวัดอะไร เครื่องเรียนว่าจะตัดสินอย่างไร และชุดข้อมูลเริ่มใหญ่ขึ้นPeople designed what to measure; machines learned how to decide; datasets grew.

  1. 2001

    Rapid Object Detection using a Boosted Cascade of Simple Features

    Paul Viola, Michael Jones · CVPR

    ทดสอบสี่เหลี่ยมมืดสว่างเล็ก ๆ นับพันแบบ คำนวณเร็วด้วย “ภาพปริพันธ์” (ภาพที่บวกสะสมไว้ล่วงหน้า ทำให้รวมพื้นที่ส่วนไหนก็ได้ทันที) แล้วเรียงต่อกันจนภาพส่วนใหญ่ถูกตัดทิ้งในไม่กี่ขั้น ทำให้กล้องถ่ายรูปทั่วไปหาใบหน้าได้ นี่คือการตรวจว่า “มีใบหน้า” ไม่ใช่การจำว่าเป็นใคร

    Thousands of tiny light-and-dark rectangle tests, made fast by an “integral image” (a picture pre-added-up, so any region can be summed instantly) and chained so most of the picture is rejected in a few steps. It put face detection into consumer cameras. Detection — that a face is there — not recognition of whose.

  2. 2004

    Distinctive Image Features from Scale-Invariant Keypoints

    David G. Lowe · International Journal of Computer Vision (first presented at ICCV 1999)

    SIFT: หาจุดที่ยังจำได้แม้ภาพจะถูกย่อ หมุน หรือเปลี่ยนแสง แล้วอธิบายแต่ละจุดด้วยฮิสโทแกรมทิศทางของความชัน การต่อภาพพาโนรามาและการจับคู่วัตถุอาศัยวิธีนี้อยู่นับสิบปี

    SIFT: find points that stay recognisable when the picture is scaled, rotated or relit, and describe each with a histogram of gradient directions. Panorama stitching and object matching ran on it for a decade.

  3. 2005

    Histograms of Oriented Gradients for Human Detection

    Navneet Dalal, Bill Triggs · CVPR

    HOG: อธิบายหน้าต่างภาพด้วยทิศทางของขอบ แล้วให้ตัวจำแนกเชิงเส้น (กฎง่าย ๆ ที่ชั่งน้ำหนักแล้วตัดสิน) ตัดสินว่า “คนหรือไม่ใช่” เป็นตัวตรวจจับคนเดินถนนมาตรฐานอยู่หลายปี

    HOG: describe a window by which way its edges point, then let a linear classifier (a simple rule that weighs things up and decides) decide “person or not”. The standard pedestrian detector for years.

  4. 2009

    ImageNet: A Large-Scale Hierarchical Image Database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, Li Fei-Fei · CVPR ใช้ในเว็บนี้Used here

    ภาพหลายล้านภาพ จัดหมวดตามโครงสร้างคำของ WordNet (พจนานุกรมอังกฤษที่จัดคำเป็นเครือข่ายความหมาย) โดยแรงงานออนไลน์ การแข่งขัน 1,000 หมวดของมันกลายเป็นสนามที่เปลี่ยนทั้งวงการ MobileNetV2 ของเราเรียนรู้การมองจากชุดนี้

    Millions of photos sorted into WordNet's categories (an English dictionary arranged as a web of word meanings) by crowd workers. Its 1,000-class challenge became the race that changed the field. Our MobileNetV2 learned to see from it.

  5. 2010

    The PASCAL Visual Object Classes (VOC) Challenge

    Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, Andrew Zisserman · International Journal of Computer Vision

    วัตถุยี่สิบประเภท กรอบที่คนวาดเอง และวิธีให้คะแนนตัวตรวจจับที่ทุกคนยอมรับ ทำให้ “การตรวจจับ” วัดผลเทียบกันได้

    Twenty object classes, boxes drawn by hand, and an agreed way to score detectors: the benchmark that made detection measurable.

03 · 2012–2018

ลักษณะที่เครื่องเรียนเองFeatures learned

ข้อมูลมากพอ การ์ดจอแรงพอ และเครื่องเริ่มออกแบบสิ่งที่ตัวเองวัดเองEnough data, enough graphics cards, and the machine began designing its own measurements.

  1. 2012

    ImageNet Classification with Deep Convolutional Neural Networks

    Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · NeurIPS (then NIPS)

    AlexNet: โครงข่ายคอนโวลูชันลึกที่ฝึกบนการ์ดจอสองใบ ชนะการแข่งขัน ImageNet แบบทิ้งห่าง ภายในไม่กี่ปี ระบบการมองเห็นเกือบทั้งหมดเปลี่ยนมาให้เครื่องเรียนลักษณะเอง

    AlexNet: a deep convolutional network trained on two graphics cards won the ImageNet challenge by a wide margin. Within a few years nearly every vision system learned its features instead of having them designed.

  2. 2014

    Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation

    Ross Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik · CVPR

    R-CNN: เสนอบริเวณที่น่าสนใจ ส่งแต่ละบริเวณเข้าโครงข่าย แล้วจำแนก ช้า แต่พิสูจน์ว่าลักษณะที่เรียนเองชนะในงานตรวจจับด้วย

    R-CNN: propose regions, run a network on each, classify. Slow — but it showed learned features win at detection too.

  3. 2014

    Microsoft COCO: Common Objects in Context

    Tsung-Yi Lin et al. · ECCV ใช้ในเว็บนี้Used here

    ภาพชีวิตประจำวันที่มีวัตถุหลายชิ้น แต่ละชิ้นถูกวาดขอบด้วยมือ ใน 80 ประเภท ตัวตรวจจับของเราฝึกจากชุดนี้ มันจึงรู้จัก “คน” “รถยนต์” “มอเตอร์ไซค์” และ “เรือ” แต่ไม่รู้จัก “ตุ๊กตุ๊ก”

    Everyday scenes with many objects, each outlined by hand, in 80 categories. Our detector was trained on it, which is why it knows “person”, “car”, “motorcycle” and “boat” — and not “tuk-tuk”.

  4. 2015

    Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

    Shaoqing Ren, Kaiming He, Ross Girshick, Jian Sun · NeurIPS (then NIPS)

    ให้โครงข่ายเสนอบริเวณเอง ตัวตรวจจับสองขั้นที่แม่นยำ และเป็นมาตรฐานอ้างอิงอยู่หลายปี

    The network proposes its own regions. Accurate two-stage detection, and the reference point for years.

  5. 2015

    U-Net: Convolutional Networks for Biomedical Image Segmentation

    Olaf Ronneberger, Philipp Fischer, Thomas Brox · MICCAI

    ติดป้ายทุกพิกเซลจากภาพฝึกไม่กี่ภาพ เป็นรูปแบบเบื้องหลังงานแบ่งส่วนภาพทางการแพทย์จำนวนมากนับจากนั้น

    Label every pixel, from few training images: the shape behind much of medical image segmentation since.

  6. 2016

    Deep Residual Learning for Image Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · CVPR

    ResNet: ให้แต่ละบล็อกเรียนแค่ “ส่วนที่ต้องแก้” จากข้อมูลที่เข้ามา แล้วโครงข่ายร้อยกว่าชั้นก็ฝึกได้ ทางลัดแบบนี้อยู่ทุกที่ในปัจจุบัน รวมถึงใน MobileNetV2

    ResNet: let each block learn only a correction to its input, and networks of over a hundred layers train. Skip connections are now everywhere — including inside MobileNetV2.

  7. 2016

    SSD: Single Shot MultiBox Detector

    Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, Alexander C. Berg · ECCV ใช้ในเว็บนี้Used here

    มองครั้งเดียว: ตารางกรอบตั้งต้นหลายขนาด แต่ละกรอบได้คะแนนทุกประเภท ไม่มีขั้นเสนอบริเวณแยก จึงเร็ว เป็นส่วนหัวของตัวตรวจจับของเรา

    One look: a grid of default boxes at several scales, each scored for every class. No separate proposal step, so it is fast. The head of our detector.

  8. 2016

    You Only Look Once: Unified, Real-Time Object Detection

    Joseph Redmon, Santosh Divvala, Ross Girshick, Ali Farhadi · CVPR

    YOLO: มองการตรวจจับเป็นการคำนวณครั้งเดียวจากพิกเซลไปเป็นกรอบและประเภท ทันเวลาจริง ตระกูล YOLO ยังพัฒนาต่อมาอีกหลายรุ่นโดยหลายทีม

    YOLO: detection as a single pass from pixels to boxes and classes, in real time. The YOLO family has continued through many versions and many authors since.

  9. 2017

    Mask R-CNN

    Kaiming He, Georgia Gkioxari, Piotr Dollár, Ross Girshick · ICCV

    เพิ่มกิ่งที่วาดขอบวัตถุแต่ละชิ้นทีละพิกเซล: การแบ่งส่วนรายวัตถุ

    Add a branch that outlines each object, pixel by pixel: instance segmentation.

  10. 2018

    MobileNetV2: Inverted Residuals and Linear Bottlenecks

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, Liang-Chieh Chen · CVPR ใช้ในเว็บนี้Used here

    โครงข่ายที่ออกแบบมาเพื่อโทรศัพท์ คอนโวลูชันราคาถูกแบบ depthwise (ทำทีละช่องสี แทนทั้งหมดพร้อมกัน) กับชั้นบาง ๆ ที่บีบตัวเลขให้น้อยลง คือเหตุผลที่เว็บนี้รันตัวตรวจจับในเบราว์เซอร์ของคุณได้ งานเดียวกันนี้ยังเสนอ SSDLite ส่วนหัว SSD แบบเบาที่ตัวตรวจจับของเราใช้

    A network designed for phones: cheap depthwise convolutions (one colour channel at a time, not all at once) and thin bottlenecks (layers that squeeze the numbers down). It is why this site can run a detector in your browser. The same paper introduces SSDLite, the lighter SSD head our detector uses.

04 · 2021–

แบบจำลองเดียว หลายงานOne model, many tasks

ข้อมูลระดับอินเทอร์เน็ต และแบบจำลองที่ถูกถามเรื่องที่ไม่เคยถูกสอนได้Internet-scale data, and models that can be asked about things nobody taught them.

  1. 2021

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

    Alexey Dosovitskiy et al. · ICLR

    ViT: ตัดภาพเป็นชิ้นสี่เหลี่ยม ปฏิบัติต่อแต่ละชิ้นเหมือนคำ แล้วใช้ทรานส์ฟอร์มเมอร์ (การออกแบบที่เทียบทุกชิ้นส่วนของภาพกับทุกชิ้นส่วนอื่น) ไม่มีคอนโวลูชันเลย และเมื่อข้อมูลมากพอก็ทำได้เทียบเท่าหรือดีกว่า

    ViT: cut the picture into patches, treat them like words, use a transformer (a design that compares every part of the picture with every other part). No convolutions at all — and given enough data, it matches or beats them.

  2. 2021

    Learning Transferable Visual Models From Natural Language Supervision

    Alec Radford et al. · ICML

    CLIP: เรียนจากคู่ภาพกับคำบรรยายจากเว็บหลายร้อยล้านคู่ ถามถึงหมวดที่ไม่เคยมีใครสอนได้ และรับเอาทุกอย่างที่เว็บพูดถึงโลกมาด้วย

    CLIP: learn from hundreds of millions of image–caption pairs from the web. It can be asked about categories nobody trained it on — and it inherits whatever the web says about the world.

  3. 2023

    Segment Anything

    Alexander Kirillov et al. · ICCV

    SAM: แบบจำลองที่วาดขอบวัตถุอะไรก็ได้ที่คุณชี้ ฝึกจากขอบวัตถุกว่าพันล้านชิ้น ซึ่งส่วนหนึ่งมันช่วยสร้างขึ้นเอง

    SAM: a model that outlines any object you point at, trained on over a billion masks — many of which it helped produce.

05

สิ่งที่เว็บนี้ใช้What this site runs

ที่ไหนWhereวิธีMethodงานต้นทางFromฝึกจากTrained on
เลนส์ “สิ่งของ” และการนับรถ นับคนThe Objects lens; every count of cars and peopleSSDLite + MobileNetV2Liu et al. 2016; Sandler et al. 2018COCO (Lin et al. 2014)
ลายนิ้วมือตัวเลขในห้องฝึกFingerprints in the training roomMobileNetV2 1.0 · 224Sandler et al. 2018ImageNet (Deng et al. 2009)
เลนส์ “ขอบ”The Edges lensตัวดำเนินการ SobelSobel operatorSobel และ Feldman, 1968Sobel and Feldman, 1968—
ค่าขีดแบ่งอัตโนมัติAutomatic thresholdวิธีของ OtsuOtsu's methodOtsu 1979—
เลนส์ “ขยับ”The Motion lensลบภาพต่อภาพFrame differencingเทคนิคพื้นฐานA classic technique—

ชุดข้อมูลทั้งสองไม่ได้ถ่ายในประเทศไทย ภาพใน COCO มาจาก Flickr ถนนไทย รถไทย และฝนเมืองร้อนจึงไม่ใช่สิ่งที่แบบจำลองเห็นบ่อยที่สุดตอนฝึก นี่คือเหตุผลหนึ่งที่เราแสดงความมั่นใจทุกครั้ง ดูขนาดไฟล์และที่มาของน้ำหนักได้ที่หน้าระบบ

Neither dataset was photographed in Thailand; COCO's pictures come from Flickr. Thai roads, Thai vehicles and tropical rain are not what these models saw most while training — one reason we always show their confidence. Sizes and the source of the weights are on the system page.

06

สิ่งที่เรายังตอบไม่ได้What we still cannot answer

ความทนทานRobustness

โครงข่ายที่ชนะคนในชุดทดสอบ อาจพังเพราะฝน ภาพเบลอ หรือสติกเกอร์แผ่นเดียว การเปลี่ยนพิกเซลเพียงเล็กน้อยจนตามองไม่เห็นก็พลิกคำตอบได้ (Szegedy et al., ICLR 2014; Goodfellow, Shlens, Szegedy, ICLR 2015) สัญญาณรบกวน ความเบลอ หมอก และการบีบอัดภาพยังทำให้ความแม่นยำลดลงมาก (Hendrycks, Dietterich, ICLR 2019) และโครงข่ายที่ฝึกบน ImageNet พึ่งพื้นผิวมากกว่ารูปร่าง (Geirhos et al., ICLR 2019) ลองดูตัวตรวจจับของเรากับกล้องตอนฝนตกกลางคืน

Networks that beat people on a test set can fail on rain, blur or a single sticker. Pixel changes too small to see can flip an answer (Szegedy et al., ICLR 2014; Goodfellow, Shlens, Szegedy, ICLR 2015). Noise, blur, fog and compression still cost a lot of accuracy (Hendrycks and Dietterich, ICLR 2019). And ImageNet-trained networks lean on texture more than shape (Geirhos et al., ICLR 2019). Watch our detector on a rainy camera at night.

อคติBias

แบบจำลองรู้จักโลกเท่าที่ข้อมูลแสดงให้มันเห็น Buolamwini กับ Gebru พบว่าระบบจำแนกเพศเชิงพาณิชย์แม่นยำน้อยกว่ามากกับผู้หญิงผิวเข้มเมื่อเทียบกับผู้ชายผิวขาว (Gender Shades, FAT* 2018) ชุดข้อมูลก็พกการตัดสินใจของคนสร้างมาด้วย หมวด “คน” ของ ImageNet ต้องถูกกรองใหม่หลายปีต่อมา (Yang et al., FAT* 2020)

A model knows the world its data showed it. Buolamwini and Gebru found commercial gender classifiers far less accurate for darker-skinned women than for lighter-skinned men (Gender Shades, FAT* 2018). Datasets carry their makers' choices: ImageNet's “person” categories had to be filtered years later (Yang et al., FAT* 2020).

ความเป็นส่วนตัวPrivacy

กล้องในที่สาธารณะเห็นคนที่ไม่ได้ยินยอมให้ศึกษา มีเทคนิคช่วยอยู่ เช่น เบลอภาพ นับโดยไม่เก็บ หรือประมวลผลบนเครื่องของผู้ใช้ แต่ยังไม่มีข้อตกลงร่วมกันว่า “มีประโยชน์” จบตรงไหนและ “การสอดแนม” เริ่มตรงไหน และกฎหมายแต่ละประเทศก็ต่างกัน คำตอบของเราอยู่ในการออกแบบ: นับ ไม่ระบุตัว ประมวลผลบนเครื่องคุณ และไม่เก็บอะไร ดูระบบและข้อกฎหมาย

Cameras in public places see people who never agreed to be studied. Techniques help — blurring, counting without storing, running on the user's device — but there is no agreement on where useful ends and surveillance begins, and laws differ by country. Our answer is in the design: count, never identify; process on your device; keep nothing. See the system and the fine print.

พลังงานEnergy

การฝึกแบบจำลองที่ใหญ่ที่สุดใช้การคำนวณมหาศาล และแนวโน้มคือใช้มากขึ้นเรื่อย ๆ (Schwartz, Dodge, Smith, Etzioni, “Green AI”, Communications of the ACM, 2020; Strubell, Ganesh, McCallum, ACL 2019 สำหรับแบบจำลองภาษา) ทางหนึ่งคือใช้แบบจำลองเล็กบนเครื่องที่คุณมีอยู่แล้ว ตัวตรวจจับของเราหนักราว 18 MB และรันบนชิปกราฟิกของอุปกรณ์คุณเมื่อทำได้

Training the largest models takes enormous computation, and the trend has been to use more (Schwartz, Dodge, Smith, Etzioni, “Green AI”, Communications of the ACM, 2020; Strubell, Ganesh, McCallum, ACL 2019, for language models). One answer is small models on devices people already own: our detector is about 18 MB and runs on your device's graphics chip when it can.

การวัดผลEvaluation

คะแนนบนชุดทดสอบที่มีชื่อเสียงไม่เท่ากับการใช้งานได้จริง เมื่อสร้างชุดทดสอบ ImageNet ใหม่ด้วยวิธีเดิม ความแม่นยำของแบบจำลองที่ทดสอบลดลงหลายจุด (Recht, Roelofs, Schmidt, Shankar, ICML 2019) และชุดทดสอบเองก็มีป้ายผิด (Northcutt, Athalye, Mueller, NeurIPS 2021 Datasets and Benchmarks) เราจึงแสดงความมั่นใจทุกคำตอบ และไม่เคยถือว่า “ตรวจไม่พบ” คือ “ไม่มี”

A score on a famous test set is not the same as working in the world. Rebuilding ImageNet's test set the same way lowered the accuracy of the models tested by several points (Recht, Roelofs, Schmidt, Shankar, ICML 2019), and test sets contain wrong labels too (Northcutt, Athalye, Mueller, NeurIPS 2021 Datasets and Benchmarks). So we show confidence on every answer, and never treat “not detected” as “not there”.