VISION
คู่มือ · อ่านเองทีละบทHandbook · read it on your own

08 · Limits and ethics · ข้อจำกัดและจริยธรรม

อ่านเองได้ทีละบท ทุกบทมาพร้อมแบบฝึกหัดให้ลองทำเองและคำถามท้ายบท เมื่ออยากรู้ให้ลึกกว่านี้ กลับไปที่ห้องเรียนบนเว็บเพื่อทดลองของจริง

Read a chapter at a time. Each comes with an exercise you can try on your own and questions to check yourself. When you want to go deeper, go back to the room on this site and run it for real.

ในหน้านี้On this page
  1. 1. The honesty rule
  2. 2. Where it fails: a field guide
  3. 3. The first limit is plumbing, not AI
  4. 4. Bias: who the errors fall on
  5. 5. Privacy by architecture
  6. 6. Thailand's PDPA, as it bears on computer vision
  7. 7. Questions to ask any vision system (a checklist)
  8. Try it yourself · ลองทำเอง
  9. Check yourself
In one sentence: a vision system's answers are guesses about pictures, not facts about the world; it fails most where conditions are worst; and what it is allowed to see is decided by design choices made long before any model runs.

← 07 · Handbook · Site: /learn, chapter 8 · /legal · Run: node examples/08-catalogue.mjs


1. The honesty rule

This project inherits one rule from FloodDash, and every room follows it:

A positive finding is a claim about the image. A negative finding would be a claim about the world — and we do not publish those.

"A car, 71%" says: in these pixels, a pattern resembling a car scored 0.71. Publishable. "No cars" would say: the road is empty. But chapters 04 and 06 showed that a jammed road, a fogged lens, a frozen feed, a car too small to resolve and a car the model has never seen all produce the same silence. So the site only ever says "nothing detected above N%" — and the readouts explain why.

2. Where it fails: a field guide

ConditionWhat happensChapter
Small or distant objectssqueezed to a few pixels at 300 × 300; recall 37% vs 94% for large ones in the simulation01, 06
Night, low lightsensor noise rises, contrast and colour vanish; detector confidence collapses02
Fog, rain on the lens, glareeverything washes toward one gray; thresholds and edges lose their grip02, 03
Crowds, traffic jamsboxes overlap; NMS merges neighbours; counts fall06
Unusual objectsa tuk-tuk, a song-thaew, a cart: COCO has no class for them, so they become "car", "truck" or nothing06
Things that look like other thingslamp post → "person 34%"; puddle → "boat"06
Wrong lesson learned"dark = river": 100% on training, 0% when lighting flips07
Deliberate attackadversarial stickers and patterns can make objects appear or vanish05
No pixels at alla third of listed cameras cannot be read by any machine herebelow

On /learn chapter 8 you can lower the resolution of a live camera step by step and watch the detector's answers degrade; on /games, "fool the machine" lets you find failures with your own camera.

3. The first limit is plumbing, not AI

Before any model runs, a city-scale system must be able to see. A sample run of node examples/08-catalogue.mjs (the counts move by tens a day; run it for today's):

video     53  live video from a host that allows reading (CORS)
still   2675  a still JPEG without CORS — relayed through server memory
view     857  viewable only on the owner's own page — no machine here can read its pixels
off      477  the owner's stream was found down by FloodDash's health probe

67% of listed cameras are machine-readable IN PRINCIPLE. In practice fewer: on
one sample, only 22 of 40 cctv.maholan.net stills answered (the rest 502).

Of ~4,000 public cameras, one in three offers no readable pixels; many of the rest answer intermittently; about 50 offer live video a browser may analyse. Coverage, uptime and permission bound what any vision system can know — long before accuracy does. When someone presents a "city-wide AI camera network", the first questions are: how many cameras, readable how often, with whose permission?

4. Bias: who the errors fall on

Errors are never evenly spread. The best-known measurement, Gender Shades (Buolamwini & Gebru, 2018), tested commercial gender classifiers on faces: error rates were as low as 0.8% for lighter-skinned men and as high as 34.7% for darker-skinned women. The models were not "biased" by intent; their training data and benchmarks under-represented some people, and nobody had measured.

For road cameras the same mechanism applies to places, not just people:

What to do about it: measure accuracy separately for the conditions and places that matter (night, rain, provinces, vehicle types), publish those numbers, and keep a human in the loop where errors cause harm.

5. Privacy by architecture

A promise in a policy can be broken quietly. An architecture that cannot do the thing is a stronger promise. This site's choices (detail in system/architecture.md):

Where pictures go
We do not…Because the design…
store camera imageryrelays stills from memory (≤ 30 s, never disk); live video never touches our server
receive your webcam or photosruns every model in your browser; there is no upload endpoint
recognise faces or read number platesships no such model; the only "person" output is a COCO class count
track individuals across frames or cameraskeeps no identities, no history, no cross-camera linkage
profile visitorshas no accounts, cookies or analytics; IP addresses live briefly in memory for rate limits only

What we deliberately refused to build, even though it would be easy: face recognition, licence-plate reading, person re-identification, a "find this person on every camera" search, and a "me / not me" training preset.

6. Thailand's PDPA, as it bears on computer vision

Not legal advice. The site's own position is on /legal.

The Personal Data Protection Act B.E. 2562 (2019) came fully into force on 1 June 2022. Questions any vision system in Thailand must answer:

SectionWhat it says (summarised)Why it matters for CV
s.6"personal data" is information that identifies a person directly or indirectlya clearly visible face is personal data even if no name is known
s.19, s.24processing needs consent or another lawful basis"the camera is public" is not by itself a lawful basis
s.25collecting from a source other than the person is restricteda camera feed is exactly that: we are downstream of everyone in frame
s.26sensitive data, including biometric data, needs explicit consent or a narrow exceptionface recognition processes biometric data — one concrete reason this site does not do it
s.4lists activities the Act does not apply to (e.g. some public-benefit, media and research uses), while still requiring securityexemptions are narrow and conditional; read them with a lawyer

Beyond the PDPA: the Criminal Code's s.309/1 (repeatedly watching or tracking a person so as to disturb their ordinary life) reaches users of camera systems, and the Computer-Related Crime Act covers scraping and attacks. This is why /legal asks every visitor not to identify, follow or profile anyone.

7. Questions to ask any vision system (a checklist)

For a city, an agency, or anyone being sold "AI cameras":

  1. What exactly does it output — boxes, counts, identities, alerts? Who sees them?
  2. What is its measured precision and recall, in our conditions: night, rain, our vehicles, our provinces? Who measured it?
  3. What does "nothing detected" trigger? Anything that treats silence as safety is dangerous.
  4. Where do pictures go, how long are they kept, who can access them? Is that enforced by architecture or by policy?
  5. Does it process biometric data (faces, gait)? Under what lawful basis?
  6. Who reviews its alerts before action is taken? What is the appeal route for a person affected?
  7. How is drift detected — dirty lenses, new vehicles, seasonal change?
  8. Can it be switched off per camera, and can a camera owner or a citizen ask for that?

Try it yourself · ลองทำเอง

เป้าหมาย · Goal: พิสูจน์ด้วยตัวเองว่า "ไม่ถูกตรวจพบ" ไม่เท่ากับ "ไม่มีอยู่" / Prove for yourself that "not detected" is not "not there".

ขั้นตอน · Steps

  1. เปิด /learn บท 8 โหลดเครือข่าย ชี้หาวัตถุที่เจอได้ก่อน แล้วลด "จำนวนพิกเซลตามแนวนอน" ลงเรื่อย ๆ — Open /learn chapter 8, load the network, find an object it detects, then lower Pixels across step by step.
  2. แสดงของที่เครือข่ายไม่รู้จัก เช่นของทำมือ ของแปลก ๆ แล้วดูว่าไม่มีกรอบ ไม่มีคำเตือน ไม่มีอะไรบอกว่า "ไม่แน่ใจ" — Show it something it has never learned: no box, no warning, nothing says "unsure".
  3. อ่าน /legal ตอนที่ว่าด้วยการไม่ระบุตัวบุคคล แล้วนึกถึงถนนของคุณเอง: เครื่องมือชุดนี้ถูกฝึกมาจากภาพแบบไหน — Read the part of /legal about never identifying people, then think about your own street: what was this trained on?

ควรเห็น · You should see

ถ้าไม่เห็น · If you do not — ถ้ายังตรวจพบอยู่ที่ความละเอียดต่ำ ให้ลดลงไปอีก ที่ 4 หรือ 2 ช่อง ไม่มีรูปร่างเหลือให้กรอบลอยจับ — If it still detects at low resolution, go lower: at 4 or 2 pixels across there is no shape left for an anchor box to catch.

Check yourself

1. A system reports "0 people in the flood zone" at 2 am. What can you conclude?

Only that no region scored above the threshold as "person" in the frames analysed. At night, at distance, in rain, recall is at its worst. It is not evidence that nobody is there — and must never be used to stand down a rescue.

2. Why is "the camera is already public" not enough to justify face recognition on it?

Because recognition creates new, sensitive data (biometric identity) that the original public view did not: it turns "someone walked past" into "this named person was here at 14:02". Under the PDPA biometric data is sensitive (s.26); ethically, it enables tracking that a passer-by never consented to.

3. Name one design choice on this site that protects privacy without relying on anyone's good behaviour.

Any of: running models in the browser (no upload path exists); relaying stills from memory only; shipping no face or plate model; no accounts or analytics.


สรุปภาษาไทย

กฎความซื่อตรง: สิ่งที่เครื่อง "เห็น" คือข้อสรุปเกี่ยวกับภาพ เผยแพร่ได้ แต่ "ไม่เห็น" จะกลายเป็นข้อสรุปเกี่ยวกับโลก ซึ่งเราไม่เผยแพร่ เพราะรถติดนิ่ง เลนส์ฝ้า กล้องค้าง หรือรถที่เล็กเกินไป ล้วนให้ความเงียบแบบเดียวกัน

จุดที่ระบบพลาด: วัตถุเล็กหรือไกล กลางคืน หมอก ฝนบนเลนส์ ฝูงชน และพาหนะที่ COCO ไม่รู้จักอย่าง ตุ๊กตุ๊ก หรือ สองแถว และข้อจำกัดแรกไม่ใช่ AI แต่เป็น การเข้าถึงภาพ: จากกล้องสาธารณะราว 4,000 ตัว หนึ่งในสามไม่มีภาพที่เครื่องอ่านได้เลย

อคติ: ความผิดพลาดไม่เคยกระจายเท่ากัน งาน Gender Shades (2018) พบความผิดพลาด 0.8% กับชายผิวขาว แต่สูงถึง 34.7% กับหญิงผิวเข้ม ทางแก้คือวัดความแม่นยำแยกตามสภาพและพื้นที่ที่สำคัญ และให้คนตัดสินใจขั้นสุดท้าย

ความเป็นส่วนตัวด้วยสถาปัตยกรรม: เราไม่เก็บภาพ ไม่รับภาพจากกล้องของคุณ ไม่จดจำใบหน้า ไม่อ่านทะเบียนรถ ไม่ติดตามบุคคล และตั้งใจไม่สร้างสิ่งเหล่านี้ ตาม พ.ร.บ.คุ้มครองข้อมูลส่วนบุคคล พ.ศ. 2562 ข้อมูลชีวภาพ (biometric) เช่นการจดจำใบหน้า เป็นข้อมูลอ่อนไหวตามมาตรา 26 ที่ต้องได้รับความยินยอมโดยชัดแจ้ง (เอกสารนี้ไม่ใช่คำแนะนำทางกฎหมาย)

Back to: the handbook · Next: System architecture →