Emerging Interfaces

Reconstructing Dense 3D Human Point Clouds From Body Heat, Using One Commodity Thermal Array

A physics-informed forward model recovers depth from a thermal image with no LiDAR, no radar, no depth camera — and no recognisable picture of anyone.

Human body point clouds are a useful representation for any system that needs to know where a person is and what shape they are making. Getting one currently means LiDAR (expensive), radar (sparse), or a depth camera (a camera pointed at people in their home).

TAP3D, posted 8 October 2026, reconstructs 3D human point clouds from body heat signatures, using low-cost thermal arrays — and claims to be the first system to do so.

Why a thermal array is a different proposition

A thermal array is a small grid of infrared sensors — the familiar cheap ones are 8×8 or 32×24 — that reports temperature per cell. Four properties matter here:

Cost. Commodity parts, orders of magnitude below LiDAR.

Density. The paper claims dense point cloud generation, where radar gives you a handful of returns.

Human sensitivity. A thermal array is specifically good at finding people, because people are reliably warmer than rooms. No segmentation model is needed to decide what is a person; the signal is the person.

Privacy. This is the property doing the most work. A low-resolution thermal image cannot identify anyone and cannot show you what they are doing in any recognisable sense. It is a warm blob at a position. For a sensor intended to live in a bedroom, a bathroom or a care facility, the difference between that and a camera is the difference between deployable and not — and no amount of on-device processing makes a camera feel like a thermal array to the person being sensed.

The authors describe the result as “privacy-first, fully passive human sensing.” Passive matters too: no emitter, nothing illuminating the room, no interference between units.

The hard part is depth

A thermal array gives you a 2D grid of temperatures. A point cloud needs 3D positions. The missing dimension has to come from physics, and the paper’s contribution is a physics-informed design integrating a forward thermal physics model with two modules.

Multi-primitive estimation does self-supervised joint recovery of depth and other thermal properties. The insight underneath this is that the apparent temperature at a sensor cell is not the body’s temperature — it is attenuated by distance, by the angle of the surface relative to the sensor, and by the emissivity of whatever is radiating. Those are confounds if you want temperature, and they are signal if you want geometry. A forward model that predicts the measurement from the geometry can be inverted, and because the model supplies the supervision, no ground-truth depth is required.

Geometric perspective fusion handles the two failures that would otherwise make this unusable in a real room: suppressing thermal interference — radiators, sunlight, a laptop, a warm patch where someone was sitting — and disentangling multiple people, who in a thermal image merge into one blob the moment they overlap.

The evaluation

Implemented with a single commodity thermal array sensor, with a dataset of 160K samples across 8 environments and 11 users.

Eight environments is the figure to weigh. Thermal sensing is notoriously environment-dependent — ambient temperature, reflective surfaces, heating systems and airflow all change what the sensor reads — so a method tested in one room tells you very little. Eight is a real attempt at generalisation, and 11 users is a reasonable spread of body sizes and surface temperatures.

Downstream results:

  • Fall detection — 91.46%
  • Indoor tracking — 21.86 cm mean absolute error
  • Human mesh recovery — 4.87 cm error

The mesh figure is the striking one. Under 5 cm of error on a recovered human mesh, from body heat, is far better than the sensing modality would suggest is possible, and it is what makes the claim interesting rather than merely novel.

21.86 cm tracking error is the honest limit: good enough to know which part of a room someone is in and how they are moving through it, not good enough for anything requiring precise position.

TAP3D is open-sourced.

What it is for, including the uses worth being wary of

The paper’s framing is assistive and medical — fall detection, tracking, mesh recovery — and for that application the privacy argument is genuinely strong. A fall detector that works in a bathroom is worth a great deal, and a camera in a bathroom is not an acceptable price.

For interactive and installation work the appeal is different and immediate: full-body input, in the dark, with no camera, at commodity cost. Thermal arrays do not care about lighting, which means a piece can work in a blacked-out room where every camera-based tracker fails, and they do not create a recording that an audience has to be warned about. A pose-driven installation with no privacy notice is a different kind of invitation.

The caution worth stating plainly: “cannot identify anyone” is a property of the resolution, not of the approach. A method that extracts a 4.87 cm body mesh from a low-resolution thermal image has demonstrated that considerably more information is present in that signal than the raw image suggests. The privacy argument holds for this sensor at this resolution, and it is the resolution, not the modality, that is protecting anyone.