@yacineMTB
Post
Post 1 of 5
@astridwilde1 you can get global shutter cameras to do more fps by simply scanning less rows. i am writing my own drivers
Post 2 of 5
60fps depth: ) these are pixels streamed over the air over wifi UDP, using fast foundation stereo i need to get end to end latency down to 20ms. inference is 10 right now. wonder if its possible (it is)
Video thumbnail from X post Watch video
Post 3 of 5
@astridwilde1 https://x.com/yacineMTB/status/2092686933759594899 i'm going to get a drone to do a backflip in real life before labour day Quoted post by kache (@yacineMTB) @shotpianist i am hacking the allwinner wifi driver so that it doesn't ack over wifi, so that it actually is true UDP, and i'm using a custom video encoding scheme meant for ultra low latency Open quoted post on X
Post 4 of 5
@astridwilde1 you can train a policy off of depth pixels by writing your own rendering engine, and you can train it in 2 minutes if you really know what you are doing
Post 5 of 5
@astridwilde1 witness me
Explanation
This is Yacine/Kache showing the pieces of a deliberately stripped-down autonomous-drone stack he is building almost from scratch. The goal is not merely “a drone with depth perception”; it is to make the entire perception→control loop absurdly fast, then train a controller in simulation and transfer it to a real drone. His concrete boast is that he’ll have the physical drone perform a backflip before Labour Day. The image is a live 3-D/depth reconstruction of the room, demonstrating that the perception side is already working.
The unusual part is how far down the stack he is optimizing. A global-shutter camera exposes the whole image simultaneously, which is particularly useful on a violently moving drone because you avoid rolling-shutter geometric distortion. But the sensor still has to read the pixels out. If you request fewer image rows—a shorter image—you move less data and can potentially run the sensor at a higher frame rate. So when he says he is “writing my own drivers” and scanning fewer rows, he means bypassing conservative vendor camera software and configuring the sensor around exactly the region/resolution his controller needs.
“60fps depth” means he is taking a stereo pair of ordinary camera images and computing the distance to surfaces about 60 times per second. Fast Foundation Stereo is NVIDIA’s 2026 real-time stereo model: given rectified left/right images, it estimates disparity, from which depth follows geometrically. It was specifically compressed and optimized for low-latency robotics. ([NVIDIA Docs][1]) His quoted “inference is 10 [ms]” means the neural depth computation itself takes ~10 ms; at 60 fps a new frame arrives every 16.7 ms, so that is already fast enough to keep up.
The harder target is “20 ms end-to-end.” That means roughly: photons hit sensors → image readout → transmit images → decode/process → infer depth → feed the control policy → produce a motor command. A 10 ms neural network does not automatically give you a 10 ms control loop: camera buffering, Wi-Fi, kernel queues, encoding, synchronization and control software can easily add tens of milliseconds.
His Wi-Fi comment is especially non-obvious. UDP itself has no acknowledgement/retransmission mechanism, but ordinary Wi-Fi usually does: at the 802.11 MAC layer, unicast frames are acknowledged and lost frames may be retransmitted. So “UDP” can still acquire variable latency underneath. He says he is hacking an Allwinner Wi-Fi driver to suppress those acknowledgements/retries. The trade is exactly what you want for aggressive real-time control: prefer a fresh packet with occasional loss over an old packet delivered reliably. A depth frame arriving 50 ms late may be worse than a missing frame.
The custom video codec serves the same philosophy. Normal video codecs are optimized primarily for bandwidth and visual quality; they may buffer frames and exploit dependencies between frames. For robot control, predictable microseconds/milliseconds can matter more than compression efficiency.
Finally, “train a policy off depth pixels by writing your own rendering engine” is the reinforcement-learning piece. Instead of collecting enormous amounts of real drone footage, he can simulate a drone and render the same sort of depth image the real stereo system produces, train a small control policy at huge simulated speed, then run that policy on the physical drone. His “train it in 2 minutes” claim is plausible only in the narrow sense implied here: a compact policy for a tightly specified maneuver, using an extremely fast custom simulator—not training some general-purpose robot intelligence.
So the whole thread is basically: remove latency and abstraction at every layer—camera, Wi-Fi, codec, depth model, simulator, controller—until vision-based learned control becomes fast enough for an acrobatic real-world drone.
[1]: https://docs.nvidia.com/tao/tao-toolkit/latest/text/cv_finetuning/pytorch/depth_estimation/fast_foundation_stereo.html?utm_source=chatgpt.com "Fast Foundation Stereo — Tao Toolkit"