l8r Twats Library

@teortaxesTex

Post

DeepSeek Vision has to improve a lot before it's usable

Image from X post

Image from X post

Image from X post

Image from X post

Explanation

What it says Teortaxes argues that DeepSeek’s vision/reasoning is still far from usable. The example asks a model to infer a cat’s mass from the “curvature of space” visible in a photo—a joke implying the cat is massive enough to deform spacetime. Instead of engaging with the implied physics, the model first treats the pavement depression as an ordinary hole and estimates a normal cat at ~4.5 kg.

Context After being challenged to “estimate the gravitational pull,” the model’s reasoning becomes a long exploratory chain: neutron-star density, Schwarzschild radius, concrete deformation, gravitational acceleration, and possible masses from ~10¹⁴ kg up to stellar scales. But it never establishes a coherent physical model linking the photograph’s geometry to spacetime curvature. Much of the calculation is therefore arbitrary rather than inference from the image.

Why it matters The post is less about image recognition than multimodal reasoning discipline. The system correctly sees the cat and pavement, but fails to infer the intended abstraction, then compensates with verbose speculative physics. That is a useful failure mode to watch: vision models can accurately describe pixels yet still be poor at identifying what quantity the user wants modeled, choosing defensible assumptions, and distinguishing calculable information from an underdetermined joke.