MISSION 001 · VISDRONE-VID 26.4059°N 92.2335°E ALT 20 M · FOV 60° · NADIR SIH 2026 · HARDWARE EDITION
AUTONOMOUS RESCUE & ENVIRONMENTAL INTELLIGENCE SYSTEM

7,081 BOXES. 23 PEOPLE. ONE ANSWER.

A detector over a disaster zone returns thousands of boxes. A rescue coordinator needs how many, where, and who first. ARES is everything in between — running on the aircraft, so losing the network costs the video feed and nothing else.

YOLOv12s · 960 PX CONF 0.18 HIGH RECALL BYTETRACK DE-DUP PI 4 ON-DEVICE NO API KEYS
MISSION EVENT LOG · REPLAYLIVE
DEMO_CLIP · 320 FRAMES24 FPS · 1280×720
7,081 DETECTIONS 333 TRACK IDS 23 SURVIVORS RECALL 0.831 1.8 CM / PIXEL 23 M FOOTPRINT 2.5 S PERSISTENCE ZERO INVENTED DATA
§ 01 — THE HARD PART

DETECTING IS EASY.
COUNTING IS NOT.

A person walks in and out of frame six times. Six people, or one? A detector has no memory between frames — and at a threshold tuned to miss nobody, it also reports every shadow. Watch what that costs, and what removes it.

DE-DUPLICATION · SCHEMATIC DETECTING
PHASE 01 — RAW DETECTION
0
Raw detections
0
Track identities issued
0
Confirmed survivors
SCHEMATIC — NOT MODEL OUTPUT PERSISTENCE FILTER · 2.5 S OF EVIDENCE REQUIRED
0
RAW DETECTIONS · WHOLE CLIP
0
IDENTITIES THE TRACKER ISSUED
0
SURVIVORS THAT HELD 2.5 S
0%
SHORT TRACKS AT THE FRAME EDGE
§ 02 — ARCHITECTURE

FOUR STAGES.
ONE CONTRACT.

Each stage exists because the one before it produces something not yet usable. They exchange a single JSON format — which is what lets three people work independently, and what lets the demo replay a stored file instead of inferring live.

YOLOv12s
1

DETECT

Boxes and confidences at 960 px, threshold 0.18. Finds people — but has no memory between frames.

BYTETRACK
2

TRACK

Persistent IDs across frames, so the same person is never counted twice. Position is still in pixels.

NADIR
3

LOCALIZE

Pixel offset → GPS from known altitude and field of view. 1.8 cm per pixel. Now they are on a map.

WEIGHTED
4

RANK

Confidence, cluster size and hazard proximity, in a formula printed on screen. Not a black box.

DETECT TRACK LOCALIZE RANK DASHBOARD
§ 03 — THE DECISION THAT DEFINES IT

A FALSE ALARM COSTS
THIRTY SECONDS.

A missed survivor cannot be recovered. The costs are wildly asymmetric, so the confidence threshold should be too — ours is 0.18, not the conventional 0.5, read off the validation run's precision-recall curve.

DEFAULT · CONF 0.500.740
ARES · CONF 0.180.831
0RECALL1.0
+91 / 1,000

More survivors found, for every thousand present. That is the entire argument for this system's most important parameter.

THE PRICE OF THAT CHOICE

At 0.18 the detector reports anything person-shaped — a shadow, a bag, a patch of rubble — and the tracker gives each one an identity. That is why 333 identities exist for 23 people.

We measured those short tracks: mean confidence 0.42 against 0.52 for long ones, and only 7% starting or ending at the frame edge. Low confidence, mid-frame, gone in a frame or two is flicker — not someone walking out of shot.

So the filter is persistence, not a higher threshold. A track must hold for 2.5 seconds. Expressed as a duration, because that is 60 frames on this clip and 3 on a Raspberry Pi.

§ 04 — HONEST SCOPE

WHAT IS BUILT,
AND WHAT IS NOT.

Nothing in the "described" row is simulated, mocked or implied anywhere in the product. We would rather answer hard questions about what is on the screen.

CAPABILITYSTATUSDETAIL
On-device inferenceBuiltYOLOv12s, ONNX-exported, Raspberry Pi 4 target
Counting & trackingBuiltByteTrack with a measured persistence filter
Geo-tagged mappingBuiltNadir projection from assumed altitude and FOV
Priority rankingBuiltA transparent weighted formula, printed on screen
Command dashboardBuiltWorks with the backend switched off
Offline resilienceBuiltA consequence of on-device inference, not a feature
RGB + thermal fusionPartialPublic thermal datasets, not physical hardware
Hazard classificationPartial3 of 7 classes — fire/smoke, flood, collapse
Autonomous nav · SLAMDescribedArchitecture only. Deliberately not built, and not mocked
NO INVENTED DATA

The hazard list is empty

Because the classifier does not exist yet. A fire on the map would make the ranking respond to something we made up, so the scoring drops the term and says it did.

NO INVENTED TELEMETRY

There is no battery gauge

There is no aircraft. The dashboard shows the constants it computed with instead, each tagged assumed, chosen, derived or measured.

REAL MODEL OUTPUT

Borrowed flight, our perception

No drone, so the clip is a public UAV dataset. Every box on it is our trained model's actual output over that footage.

§ 05 — EDGE DEPLOYMENT

A SEARCH UAV DOES NOT
NEED 30 FPS.

It needs to not miss the ground. At 20 m the camera sees 23 m of it, and the aircraft takes seconds to cross its own footprint — so every patch is seen in several consecutive frames even at one frame per second.

FRAME n n + 1 n + 2 n + 3 23 M OF GROUND DRONE ADVANCES SAME PERSON 4TH LOOK
WHY IT HOLDS

Redundancy is free

A miss in one frame is recovered in the next. Thirty FPS gives thirty looks at the same ground, twenty-nine redundant, at thirty times the power budget.

WHY IT IS SLOW

Attention on a Cortex-A72

YOLOv12 is attention-centric and trades speed for accuracy. On a Pi 4 with no accelerator that trade is real, and worth stating rather than hiding.

WHAT COMES NEXT

Quantisation, then NCNN

INT8 is roughly 4× smaller and much faster on CPU. NCNN is hand-optimised for ARM. Both are measurements we have not taken yet.

§ 06 — MEASURED

THE NUMBERS.

Trained on combined C2A and VisDrone under one unified person class. Every figure from the validation run — none estimated.

0
PRECISION · OF WHAT IT CALLS A PERSON
0
RECALL · AT THE DEFAULT THRESHOLD
0
mAP50 · SMALL AERIAL TARGETS
0
mAP50-95 · STRICT BOX PRECISION

For context: YOLO12s scores 48.0 mAP50-95 on COCO — mostly large, centred, well-lit objects. ARES scores 49.4 on small humans seen from directly above.

ARES command dashboard during playback
ARES dashboard map and rescue queue
ARES video panel with detection boxes
§ 07 — TEAM

WHO BUILT IT.

DEWANG

Detection model, hazard classifier, on-device benchmark, backend and dashboard.

ROBIN

Tracking, pixel-to-GPS localization, and the rescue priority scoring logic.

UJJAINI

Pitch, presentation and the recorded demo.

§ SIGN OFF

FIND THEM FASTER.
MISS NOBODY.

SMART INDIA HACKATHON 2026 · HARDWARE EDITION