ความทนทานRobustness
โครงข่ายที่ชนะคนในชุดทดสอบ อาจพังเพราะฝน ภาพเบลอ หรือสติกเกอร์แผ่นเดียว การเปลี่ยนพิกเซลเพียงเล็กน้อยจนตามองไม่เห็นก็พลิกคำตอบได้ (Szegedy et al., ICLR 2014; Goodfellow, Shlens, Szegedy, ICLR 2015) สัญญาณรบกวน ความเบลอ หมอก และการบีบอัดภาพยังทำให้ความแม่นยำลดลงมาก (Hendrycks, Dietterich, ICLR 2019) และโครงข่ายที่ฝึกบน ImageNet พึ่งพื้นผิวมากกว่ารูปร่าง (Geirhos et al., ICLR 2019) ลองดูตัวตรวจจับของเรากับกล้องตอนฝนตกกลางคืน
Networks that beat people on a test set can fail on rain, blur or a single sticker. Pixel changes too small to see can flip an answer (Szegedy et al., ICLR 2014; Goodfellow, Shlens, Szegedy, ICLR 2015). Noise, blur, fog and compression still cost a lot of accuracy (Hendrycks and Dietterich, ICLR 2019). And ImageNet-trained networks lean on texture more than shape (Geirhos et al., ICLR 2019). Watch our detector on a rainy camera at night.
อคติBias
แบบจำลองรู้จักโลกเท่าที่ข้อมูลแสดงให้มันเห็น Buolamwini กับ Gebru พบว่าระบบจำแนกเพศเชิงพาณิชย์แม่นยำน้อยกว่ามากกับผู้หญิงผิวเข้มเมื่อเทียบกับผู้ชายผิวขาว (Gender Shades, FAT* 2018) ชุดข้อมูลก็พกการตัดสินใจของคนสร้างมาด้วย หมวด “คน” ของ ImageNet ต้องถูกกรองใหม่หลายปีต่อมา (Yang et al., FAT* 2020)
A model knows the world its data showed it. Buolamwini and Gebru found commercial gender classifiers far less accurate for darker-skinned women than for lighter-skinned men (Gender Shades, FAT* 2018). Datasets carry their makers' choices: ImageNet's “person” categories had to be filtered years later (Yang et al., FAT* 2020).
ความเป็นส่วนตัวPrivacy
กล้องในที่สาธารณะเห็นคนที่ไม่ได้ยินยอมให้ศึกษา มีเทคนิคช่วยอยู่ เช่น เบลอภาพ นับโดยไม่เก็บ หรือประมวลผลบนเครื่องของผู้ใช้ แต่ยังไม่มีข้อตกลงร่วมกันว่า “มีประโยชน์” จบตรงไหนและ “การสอดแนม” เริ่มตรงไหน และกฎหมายแต่ละประเทศก็ต่างกัน คำตอบของเราอยู่ในการออกแบบ: นับ ไม่ระบุตัว ประมวลผลบนเครื่องคุณ และไม่เก็บอะไร ดูระบบและข้อกฎหมาย
Cameras in public places see people who never agreed to be studied. Techniques help — blurring, counting without storing, running on the user's device — but there is no agreement on where useful ends and surveillance begins, and laws differ by country. Our answer is in the design: count, never identify; process on your device; keep nothing. See the system and the fine print.
พลังงานEnergy
การฝึกแบบจำลองที่ใหญ่ที่สุดใช้การคำนวณมหาศาล และแนวโน้มคือใช้มากขึ้นเรื่อย ๆ (Schwartz, Dodge, Smith, Etzioni, “Green AI”, Communications of the ACM, 2020; Strubell, Ganesh, McCallum, ACL 2019 สำหรับแบบจำลองภาษา) ทางหนึ่งคือใช้แบบจำลองเล็กบนเครื่องที่คุณมีอยู่แล้ว ตัวตรวจจับของเราหนักราว 18 MB และรันบนชิปกราฟิกของอุปกรณ์คุณเมื่อทำได้
Training the largest models takes enormous computation, and the trend has been to use more (Schwartz, Dodge, Smith, Etzioni, “Green AI”, Communications of the ACM, 2020; Strubell, Ganesh, McCallum, ACL 2019, for language models). One answer is small models on devices people already own: our detector is about 18 MB and runs on your device's graphics chip when it can.
การวัดผลEvaluation
คะแนนบนชุดทดสอบที่มีชื่อเสียงไม่เท่ากับการใช้งานได้จริง เมื่อสร้างชุดทดสอบ ImageNet ใหม่ด้วยวิธีเดิม ความแม่นยำของแบบจำลองที่ทดสอบลดลงหลายจุด (Recht, Roelofs, Schmidt, Shankar, ICML 2019) และชุดทดสอบเองก็มีป้ายผิด (Northcutt, Athalye, Mueller, NeurIPS 2021 Datasets and Benchmarks) เราจึงแสดงความมั่นใจทุกคำตอบ และไม่เคยถือว่า “ตรวจไม่พบ” คือ “ไม่มี”
A score on a famous test set is not the same as working in the world. Rebuilding ImageNet's test set the same way lowered the accuracy of the models tested by several points (Recht, Roelofs, Schmidt, Shankar, ICML 2019), and test sets contain wrong labels too (Northcutt, Athalye, Mueller, NeurIPS 2021 Datasets and Benchmarks). So we show confidence on every answer, and never treat “not detected” as “not there”.