Problem Solving Report · Mode Switch Framework

Stop Optimising the Prior.
Design the Sweep.

The sparsity claim in freehand 3D ultrasound is exhausted because thirteen years of work has optimised the regulariser while leaving the sampling operator untouched. The unworked seam is acquisition design — and it happens to be the one a solo researcher with RF access and a phantom rig can actually win.

Produced by Drew Bruce Date 18 July 2026 Source Sparsity-Aware Freehand 3D Ultrasound Research Brief Decision Which research direction to commit the next 6–18 months to

General public explanation

What this is about, in plain English

We want ultrasound scans that take less time and less skill, without the 3D picture quietly inventing things that were never there.

From a 2D ultrasound slice to a 3D organ

One freehand sweep: each flat image is recorded in space, then rebuilt into a volume — with the parts nobody measured left honestly flagged.

1 Capture one thin slice 05 1015 Depth 16 cm Gain 62% Focus 9 cm Frame 0124 REC 2 Record where it belongs Z Y X 3 Reconstruct the volume Measured Weak coverage Z Y X Measured coverage 74% Measured slices anchor the reconstruction. Unobserved regions stay flagged — not silently invented.

An ordinary ultrasound probe sees a single thin slice of the body at a time — like a torch beam that lights up one flat sheet. To get a 3D picture, a sonographer sweeps the probe across the skin and a computer stacks all those slices back together into a solid volume.

skin surface probe One thin slice at a time swept across the skin stack them A 3D volume rebuilt from the slices

The basic idea: sweep the probe, collect flat slices, rebuild a solid 3D picture from them.

So what's the problem?

Sweeping thoroughly takes time and a steady, trained hand. Everyone would like to sweep less — fewer passes, fewer slices — and still get a trustworthy 3D picture. The trouble is that when you collect fewer slices you leave gaps, and something has to fill them in. Modern software is very good at filling gaps with something that looks right. That is exactly the danger: a smooth, convincing, professional-looking 3D image can contain anatomy nobody actually measured. In medicine, a confident-looking picture that is quietly wrong is worse than an obviously rough one.

What has everyone been trying?

For over a decade, researchers have attacked this by making the gap-filling software cleverer — better mathematics, then neural networks, then AI image generation. Drew's original research brief reached an honest and slightly deflating conclusion: that road is largely used up. The clever-software approach was already published back in 2013, and everything since has been variations on it.

The idea in this report

Everyone has been trying to guess better. Almost nobody has asked a more basic question: what if you swept differently in the first place?

Think of mowing a lawn. If you always mow in neat parallel stripes and then skip every other stripe to save time, you get long unmown strips — and no amount of clever guessing tells you what was growing there. But if you cross the lawn at varied angles, the same amount of mowing covers far more ground and leaves no long blind strips. Where you point the probe turns out to matter more than how sophisticated your gap-filler is.

blind strip nothing measured here Neat parallel sweeps
Varied, crossing sweeps

Same number of passes in both pictures. Only one of them leaves a region where nothing was ever measured.

What we are recommending, in plain terms

First, find out whether it's even worth it. Before designing anything clever, calculate the best possible result you could ever get if you were free to point the probe anywhere at all — a theoretical "perfect scan". If that perfect scan is barely better than today's neat stripes, the whole idea is dead and we've learned it in a week instead of a year. That single check is the cheapest and most important thing on the list.

Second, design sweeps a real human can actually do. There's no point inventing a pattern that only a robot could perform. The sweep has to keep the probe against the skin, move at a speed a wrist can manage, and be describable to a sonographer in one sentence.

Third, borrow a lie-detector from another field. Scientists who image viruses and proteins (cryo-EM) faced the same "is this real or did the software invent it?" problem years ago, and solved it: split your data in half, rebuild the picture twice independently, and see whether the two halves agree. If they don't, something was invented. Ultrasound has never adopted this test. It would let Drew prove a picture is measured rather than merely plausible.

Fourth, use crossing sweeps to correct the shaky hand. Whenever a sweep crosses over a place it has already been, you get to compare — and any disagreement is a direct measurement of how much the tracking has drifted. The scan effectively checks its own accuracy as it goes.

Why it matters

If this works, ultrasound scans get faster and less dependent on operator skill, while becoming more honest about what they did and didn't see — including flagging the regions a clinician shouldn't trust. That combination matters most where trained sonographers are scarce, which is most of the world.

The rest of this report is the technical version of the same argument, with the reasoning, the alternatives considered, honest scoring, and the things most likely to go wrong.

1 The Real Problem

Your brief asked: given sparse freehand slices, what reconstruction method is still novel? That question has an honest but narrow answer — it sends you into a crowded algorithm race against better-resourced groups with clinical data you don't have.

The Mode Switch turned up a different question hiding underneath. Your own brief lists "Sampling is not incoherent" as Risk #1 — a caveat to report. But compressed-sensing theory is unambiguous that recovery quality is governed by the measurement operator, not the prior. Which means your Risk #1 isn't a risk at all. It's the single unexploited variable in the entire field, misfiled as a limitation.

Original framing

"Sparse data is given. How do I reconstruct better from irregular, clustered freehand slices than TV/ADMM did in 2013?"

Redefined problem

"The sweep is a design variable, not a given. What acquisition geometry makes the missing content identifiable — and how do I prove the recovered volume is real, not plausible?"

This reframe matters because it changes what you are competing on. In the first framing you are competing on modelling and compute, where you are outgunned. In the second you are competing on problem formulation, where nobody is currently standing — and where your two genuine assets (RF access, and a controllable phantom/simulation rig) are exactly the right assets.

Prior-art check — targeted search through 18 July 2026

Incoherent or jittered sweep design as the lever in freehand 3D US: not found — the trajectory-design literature is almost entirely MRI. Gold-standard half-map FSC: a cryo-EM standard, unused in US reconstruction. Coherence factor: mature inside beamforming, never lifted to weight a freehand volume reconstruction. Conformal prediction for imaging inverse problems: becoming crowded (QUTCC, conformal risk control for CT) — use it as a component, not a headline. This is a narrative search, not a systematic review or a patentability opinion.

2 Mode Diagnosis

Your brief had already done excellent CLARIFY work — it separated the exhausted claim from the defensible one and correctly formulated the inverse problem. Then it converged: it named a recommended claim and laid out a four-stage validation programme. That is EXECUTE thinking. But the highest-value EXPLORE work was still unfinished, and the tell is precisely that Risk #1 was classified as a hazard to disclose rather than a lever to pull.

CLARIFYdone well in the brief
EXPLORE← where you needed to be
EVALUATEpartly done
EXECUTEwhere the brief landed

Converging into EXECUTE before EXPLORE is finished is the most expensive mode error there is, because it locks a validation programme around a claim you haven't finished stress-testing. Six months of forward-model engineering, correctly executed, against a claim that was never the best available one.

Overall confidence in this recommendation

74% — the diagnosis rests on well-established sampling theory and a clear gap in the prior art, but the size of the achievable gain in ultrasound specifically is unmeasured, and physical sweep constraints may compress it.

Considered estimate based on judgment and a targeted literature search — not measured data.

3 How We Got Here

The audit trail — each named method and the one insight it produced.

Method What it produced
The Mode Switch The brief had converged to EXECUTE while the most valuable EXPLORE work was unfinished.
Ask What Constraints Are Real vs Assumed "The sweep is whatever the sonographer happens to do" is an assumed constraint. Sonographers already follow taught protocols; protocols can be redesigned.
First-Principles Thinking Recovery from undersampled data is governed by the sampling operator's incoherence. The field has spent thirteen years optimising the other term.
Separate Symptoms from Root Causes "TV oversmooths and staircases" is a symptom. The cause is an unidentifiable null space produced by clustered, correlated sampling — no prior fixes that, it only papers over it.
Borrow the Adjacent Expert Nine fields raided. Every one that beat this problem beat it on the acquisition side. See section 4.
Map the System Not the Symptom Revealed a reinforcing loop that explains why field effort keeps rising while acquisition reduction stalls. See section 5.
Follow the Energy ~99% of the field's effort sits on the algorithm side of the equation. The unworked seam is acquisition design.
Look for the Simplest Viable Solution A sweep protocol is simpler than a new prior, needs no new mathematics, and travels to every existing solver — including your competitors'.
Solve for Reversibility Simulation probes are near-free to undo. Committing to a full RF forward model is not — so sequence it second.
Test Assumptions Quickly The riskiest assumption ("sampling design measurably matters in US") is testable in days, not months.
Distinguish Must-Haves from Nice-to-Haves Measured-slice consistency and an honest coverage map are must-haves. A more sophisticated prior is a nice-to-have.
Pre-Mortem the Solution Three specific ways this fails: boiling the ocean, simulation-flattered results, and a protocol no human hand can perform. See section 8.

4 The Adjacent-Field Raid

You asked to go wide and borrow from other industries. Here is the striking finding: every mature field that solved a structurally identical problem solved it by changing the acquisition, the validation, or both — not the regulariser. Ultrasound is the outlier still sampling on a smooth raster and blaming the algorithm.

Field Their version of your problem What they actually did Your borrow
Exploration seismology Shots are sparse, irregular and extremely expensive Jittered / randomised shot geometry; full-waveform inversion for medium parameters rather than images; illumination and null-space maps published as results Randomise the sweep, not just the solver. Invert for the medium. Publish the null-space map.
MRI compressed sensing Undersampled k-space Variable-density and Poisson-disc sampling design; TPSF as a quantitative incoherence metric; unrolled networks with hard data-consistency layers Produce the first real incoherence number for a freehand US operator. Hard-enforce measured-slice consistency.
Radio interferometry (VLA / EHT) Image from sparse, irregular uv-coverage Coverage reported as a first-class result; closure quantities invariant to per-station calibration error; synthetic-data challenges and multiple independent pipelines to prove nothing was invented Seek pose-error-invariant observables. Adopt their anti-hallucination protocol wholesale.
Cryo-EM Thousands of 2D slices at imperfectly known poses → one 3D volume Gold-standard half-map FSC, local resolution maps, joint pose+volume refinement with explicit overfitting controls Half-sweep FSC as a borrowed, hard-to-argue-with hallucination detector. Local resolution = your voxel confidence map.
Industrial NDT / beamforming Can I trust this reconstructed flaw enough to ground an aircraft? Coherence factor and Van Cittert–Zernike — physics-derived, non-learned per-voxel reliability computed from channel data Replaces the heuristic Wi in your objective with physics — and it requires the RF you have and most competitors don't.
Sparse-view CT / nuclear medicine Prove fewer views is diagnostically safe Recovery curves as the standard evidence format; model observers (Hotelling / channelised) as a quantitative stand-in for readers Clinically meaningful task metrics without the reader study you can't presently run.
Numerical weather prediction Where should the next observation go? Targeted / adaptive observation, ensemble sensitivity, OSSEs — they fly aircraft to where uncertainty reduction is greatest Rigorous machinery behind "uncertainty-driven next-view guidance", instead of a heuristic.
Robotics SLAM Drift accumulating along a freehand trajectory Loop closure; Fisher-information and D-optimal view selection Crossing sweeps reframed as designed drift constraints with a stated guarantee — not a nice-to-have listed under "sweep types".
Statistics / conformal prediction Guarantees without distributional assumptions Conformal risk control with finite-sample coverage guarantees Provable coverage on the "not observed" mask. Component, not headline — this lane is filling fast.
The pattern across all nine

Two levers recur, and neither is the one ultrasound has been pulling. Lever one: design the measurement. Seismic jitters its shots, MRI designs its k-space mask, astronomy exploits Earth rotation, cryo-EM depends on randomly oriented particles for Fourier coverage. Lever two: build an instrument that proves you didn't invent the answer. Cryo-EM has FSC, astronomy has closure quantities and blind challenges, CT has model observers, NDT has coherence factor. Freehand ultrasound has PSNR and SSIM — metrics that your own brief notes can improve while the geometry gets worse.

5 Why the Field Is Stuck

Mapping the system rather than the symptom exposes a loop that feeds itself. Each turn of it produces more sophisticated work and no more acquisition reduction, because the one term that governs recoverability never changes.

Smooth raster sweeps the taught, comfortable protocol Clustered, correlated samples large, unidentifiable null space Few-frame reconstructions look bad gaps, blur, staircasing Invest in stronger priors TV → implicit → diffusion Priors invent anatomy so validation burden rises, and trust falls

The sampling operator is the one node the loop never touches. This is why the field feels busy while acquisition reduction stalls — and why your brief's verdict ("the broad claim is exhausted") is correct about the loop but wrong about the problem.

Where the field's effort actually goes

Reading your brief's own prior-art table as an effort census makes the displacement obvious.

Rough allocation estimated from the five lines of work in the brief's state-of-the-art table and their relative publication volume. Illustrative judgment, not a bibliometric measurement.

6 Options Considered

Five directions you could credibly choose, including staying with the brief's plan. Scored on Likelihood of delivering your three stated wins (defensible novelty, acquisition reduction, clinical impact), Confidence in that estimate, and Reversibility.

Option A · Status quo

Execute the brief as written

Uncertainty-aware, physics-guided sparse inverse reconstruction with bounded pose refinement and calibrated failure maps, validated through the four-stage programme. Technically sound and honestly scoped.

Trade-off: buys a defensible paper at the cost of competing head-on in the most crowded lane, against groups with clinical data, compute and headcount you don't have. Its differentiator — uncertainty — is being commoditised by the conformal-prediction wave right now.

Likelihood
45%
Addresses a real gap, but competes on modelling and compute where you are outgunned, and the "uncertainty" differentiator is eroding.
Confidence
80%
High — the brief's own prior-art table is strong, specific evidence about how crowded this lane is.
Reversibility
Medium
Months of forward-model engineering before you learn whether it differentiates. Components are reusable; the time isn't.
Option B · Favoured direction, variant 1

Sampling-first: the incoherence programme

Treat the sweep as the design variable. Quantify the actual coherence and null-space structure of real freehand operators, design jittered / multi-angle / loop-closing protocols, and produce the first honest recovery curve versus sampling design for freehand US. The output is a protocol plus a theory-grounded diagnostic — algorithm-agnostic, so it travels to every existing solver.

Trade-off: gives up the "novel algorithm" identity that reviewers pattern-match to, in exchange for owning a question nobody has asked. Some reviewers will want a new method attached.

Likelihood
78%
Attacks the governing variable directly; every adjacent field shows large gains from sampling design; targeted search found no prior claim.
Confidence
62%
Moderate — the existence of the gap is well evidenced, but the achievable gain in US is unmeasured and physical sweep limits may compress it.
Reversibility
High
Pure simulation to start. A week or two tells you whether the recovery curves separate at all.
Option C · The big, tempting move

RF-native: reconstruct the medium, not the picture

Abandon the B-mode volume as the target. Estimate a view-invariant scatterer / impedance field from pulse-echo RF, with coherence-factor and Van Cittert–Zernike reliability inside the optimiser. This dissolves your brief's Risk #2 ("no single view-invariant volume") rather than mitigating it, and it is the deepest moat available to you.

Trade-off: the highest ceiling and the highest chance of an expensive null result. The adjacent successes here — full-wave inverse scattering — rely on transmission-mode ring arrays, not freehand pulse-echo. That is a large extrapolation.

Likelihood
55%
Strongest moat and best physics story, but view-invariance from freehand pulse-echo RF is genuinely unsolved and may remain so.
Confidence
40%
Low — thin transferable evidence. The successful precedents use a fundamentally different acquisition geometry.
Reversibility
Low
Deep commitment in modelling, calibration and compute before you get any signal about whether it works.
Option D · Favoured direction, variant 2

The trust instrument: borrow the validation stack

Import cryo-EM's gold-standard half-map FSC and local resolution maps, CT's model observers, and conformal risk control — and build the measurement instrument the field lacks. A methodology and benchmark contribution rather than a method contribution.

Trade-off: nearly unscoopable and useful to everyone, but on its own it delivers little acquisition reduction. It is an enabler, not a destination.

Likelihood
72%
Very likely to produce a real contribution; unlikely by itself to deliver your acquisition-reduction win.
Confidence
80%
High — FSC, local resolution and model observers are mature, battle-tested instruments. Transfer risk is genuinely low.
Reversibility
High
Small, self-contained, and valuable regardless of which direction ultimately wins.
Option E · Recommended

Couple B and D now; sequence C as paper two

Run the sampling-design programme and the borrowed validation instrument together, because each makes the other credible: sampling design is the claim, and the trust instrument is what stops a reviewer saying "your prior just made that up." Hold the RF scatterer-field work as the follow-on, opened only if the probes justify it.

Trade-off: two threads instead of one, and it defers the deepest moat by a year. In exchange, every stage is a decision point rather than a bet.

Likelihood
80%
Pairs the highest-leverage lever with the instrument that makes its claims defensible, and sequences the risky part behind evidence.
Confidence
70%
Rests on B and D, both individually well-supported. The sequencing itself adds little risk.
Reversibility
High
Staged with explicit gates. Option C stays opt-in, entered only on evidence.
Option F · Added after review

Zero-hardware multi-channel: a view-invariant scaffold from the probe you already have

Take a second channel that your existing probe already produces — Doppler flow or elastographic stiffness — and use it as an angle-invariant geometric scaffold for the view-dependent B-mode reconstruction, coupled by cross-gradient joint inversion. This attacks Risk #2 ("no single view-invariant volume") by a completely different mechanism than Option C, at a fraction of the cost. See section 9.

Trade-off: stiffness and flow channels are sparser and noisier than B-mode and are not available everywhere in the volume, so the scaffold is partial. In exchange it needs no new hardware, no calibration rig and no ethics amendment — and it is fully simulatable.

Likelihood
68%
Attacks the view-invariance risk directly with a mature, well-posed coupling term; weakened by the scaffold being partial and channel-dependent.
Confidence
55%
Cross-gradient inversion is proven in geophysics, but its transfer to speckle-dominated US channels is untested — that is the open question.
Reversibility
High
One extra term in an objective you are building anyway. Costs days, not months, and is discardable.

Likelihood and Confidence are considered estimates with stated reasoning — not measured data.

Decision matrix — options against your stated wins

Option Defensible
novelty
Acquisition
reduction
Clinical /
translational
Solo + RF +
phantom feasible
IP / moat
potential

Cell values are judgement of fit on a 0–100 scale, not measured performance. Darker = stronger fit.

The strong-and-safe quadrant is top-right: high impact, easy to undo. Estimates, not measured.

7 The Recommendation

Recommended move

Run three parallel safe-to-fail probes in the digital twin before committing to any forward model.

The breadth gate applies here: the territory is genuinely novel, and in simulation these three options are cheap and independent. That makes parallel probes the right shape, not a single wedge. Each has an explicit amplify/dampen criterion, so at the gate you are reading evidence rather than defending a preference.

Learning target: Does where we sample buy more acquisition reduction than how we regularise — and can we prove the recovered anatomy is measured rather than invented?

Mine the evidence you already have first. Before generating anything new, re-implement Afonso & Sanches 2013 (your ref 4) and a kernel-regression baseline (ref 5) on your own twin. They are the richest evidence in hand: they define the exact boundary your novelty has to clear, and you will need them as the comparison arm in every probe anyway. Do this in week 1 — not as a courtesy citation, but as your yardstick.

Probe 1 · The core claim

Does sweep geometry beat sweep density?

Write one script — call it sweep_coherence.py — that takes a trajectory and returns (a) the mutual coherence / TPSF sidelobe level of the operator H, and (b) a per-voxel null-space energy map. Then compare four trajectories at identical frame counts with the same solver: smooth linear raster (the baseline everyone uses), jittered-spacing linear, multi-angle fan, and cross-hatch with deliberate loop closures.

Watch
Recovery curve — reconstruction error versus frames retained, swept 90% → 10%, for each trajectory. Twenty seeds minimum so you can see the noise band.
Amplify if
At 50% frames, the best designed trajectory beats the smooth raster by a margin clearly exceeding the across-seed spread. That single plot is your paper.
Dampen if
Separation is under ~5% and inside the noise band across ≥20 seeds. Sampling is not the lever in US — fall back to Option D and re-open C.
Guardrail
Constrain the design space to sweeps a human hand can actually perform: probe stays in contact, angular rate within normal scanning. A physically impossible protocol makes the acquisition-reduction claim fictional.
Probe 2 · Your RF moat

Does physics-derived confidence beat heuristic weighting?

Add a coherence-factor channel to the twin's output and use it directly as the measurement-confidence term Wi in your objective. Compare three weightings: uniform, B-mode heuristic (the current norm), and CF / Van Cittert–Zernike.

Watch
Does CF weighting downweight shadowed and dropout voxels that B-mode weighting happily keeps? Plot a reliability diagram of predicted confidence against actual held-out error.
Amplify if
CF weighting improves held-out slice NLL and the confidence map correlates with the true error map. That is your calibrated failure map, derived from physics rather than learned — and it needs RF, which is your defensible asset.
Dampen if
CF adds nothing over the B-mode heuristic. Then RF is not your moat, and Option C's ceiling drops sharply.
Probe 3 · Do this one first if time is short

Can a borrowed instrument catch hallucination that PSNR and SSIM miss?

Split each acquisition into two interleaved half-sweeps, reconstruct them independently, and compute the FSC and a local-resolution map between them — cryo-EM's gold-standard protocol, transplanted. Then deliberately engineer the failure case: a strong prior that improves PSNR and SSIM while erasing a lesion or inventing structure.

Watch
Whether half-sweep FSC degrades on exactly the reconstructions where SSIM improves.
Amplify if
You get that one figure — "this reconstruction scores better and is anatomically wrong; here is the instrument that catches it." It is a paper hook, a reviewer-proof argument, and it services every other option you might take.
Dampen if
FSC tracks SSIM closely and adds no discriminating power in the US setting. Then use local-resolution maps only, and lean harder on model observers.

First sequence

Click to tick off. Roughly six weeks to a real decision gate.

8 Sharpening the Method

A formalisation of the trajectory generator followed this report — continuous pose stream θ(t), hand-achievability constraints on contact, velocity and acceleration, harmonic sweep primitives, and a "one-sentence rule" for cognitive feasibility. That is Probe 1 written up properly, and the guardrails sit exactly where they should. Pressure-testing it surfaced five refinements that change what to optimise, what to run first, and what to measure.

Refinement 1

Optimise observability, not coherence

The natural objective is to minimise the mutual coherence of the forward operator, minΩ μ(H(θ(Ω))). I'd argue against it, for three reasons that stack.

It is at war with your own constraint set. Mutual coherence is minimised by maximally spread, decorrelated measurement rows. Bounded acceleration and low-frequency harmonics force the opposite — neighbouring frames must be correlated. You are therefore not optimising toward incoherence; you are finding the least-coherent member of a deliberately coherent family, and that minimum may sit far above anything compressed sensing would call useful.

It is a weak and badly-behaved metric. μ is a maximum over column pairs — worst-case, dominated by a single pair, yielding notoriously pessimistic recovery bounds. It is non-smooth, so it optimises poorly, and forming the full Gram matrix for a 256³ volume is punishing.

It answers a question you are not asking. Coherence governs sparse recovery. Soft tissue is not genuinely sparse. Your real problem is identifiability: does the operator span the volume, or are there directions nothing measured? That is Risk #1 stated properly — and it is spectral, not combinatorial.

Mutual coherence μ one number, from the single worst pair everything else is ignored
Observability λ(HᵀWH) the whole spectrum — where the null space is λmin the blind direction

Coherence reports one worst pair. Observability reports the entire spectrum — and λmin is your brief's Risk #1 turned into a number you can optimise.

So replace min μ with an optimal experimental design objective on the normal operator:

E-optimality   maxΩ λmin(HᵀWH)   — directly shrinks the worst unobserved direction
D-optimality   maxΩ log det(HᵀWH + Σ₀⁻¹)   — equals expected information gain in the linear-Gaussian case

Both are smooth, so they optimise properly. Both estimate cheaply and stochastically — Lanczos for extremal eigenvalues, Hutchinson for the trace and log-det — without ever forming the Gram matrix. And both subsume next-view guidance as the sequential/greedy case, which folds your weather-prediction borrow into the same framework. Worth knowing that the Bayesian optimal experimental design literature uses "where to place a handheld ultrasound probe" as a standard motivating example: the machinery is recognised as applicable to your exact problem, and nobody has carried it into freehand sweep design.

Refinement 2

Run the oracle before you design anything

This is the change I'd act on soonest, because it is cheap and it can retire the entire programme in a week rather than a year.

Before optimising any human-feasible trajectory, compute the unconstrained oracle: greedy D-optimal frame selection with no kinematic constraints at all, free to choose any subset of poses. That gives you the ceiling — the most that any sampling design could ever buy. Then measure the oracle gap: what fraction of that ceiling does a constrained, hand-achievable harmonic sweep actually capture?

Illustrative shape only — the whole point of the experiment is that the true curve is unknown. Producing this plot for real is the contribution.

The read is decisive either way. If the oracle buys only a few percent over a smooth raster, sampling design is not the lever in ultrasound, the Section 10 objection was right, and you have established that rigorously and cheaply — publishable as a negative result with an instrument attached. If the oracle buys a lot and your feasible sweep captures a decent share of it, you have both a paper and a roadmap for closing the remainder.

The gap this closes: as originally formulated, the generator optimises within the feasible set without ever establishing that the feasible set contains anything worth having. The oracle establishes that first, for a fraction of the cost.

Refinement 3

The protocol priority is inverted — go elevational first

The jittered-spacing protocol modulates x(t) in-plane while pinning z = 0 and α = 0. But in-plane is the direction freehand ultrasound already samples densely and resolves well. The bottleneck — and where the null space actually lives — is elevational: out-of-plane resolution is governed by beam thickness, and that is precisely the term your brief singles out for the forward model. Jittering in-plane mostly reshuffles data you already have.

In-plane (lateral × axial) already dense — jitter here reshuffles good data fine sampling, well-resolved Elevational (out-of-plane) thick beam, sparse planes — this is where the null space lives unmeasured unmeasured pale bands = beam thickness (elevational PSF) · dark bars = measured planes

Schematic. The asymmetry between the two panels is the whole argument for reordering the protocols.

So the rocking / fan protocol — modulating the elevational angle β — should be the primary object of study, not the second-listed variant. I would also relax the flat-plane constraint. Small, controlled variation in contact pressure and standoff changes elevational sampling in a way a hand genuinely can produce; it couples into tissue deformation, but that is a real trade-off worth measuring rather than assuming away.

Refinement 4

Design for robustness — and turn loop closure into a drift sensor

The formulation assumes θ is known exactly. Your brief's own risk table says the opposite: geometry dominates, and sub-millimetre pose error creates blur that regularisation merely hides. A trajectory with excellent observability that collapses under 0.5 mm of pose error is worthless in a lab and unsafe in a clinic.

Make it a robust design — optimise the worst case over a pose-perturbation ball rather than the nominal operator:

maxΩ   min‖δθ‖ ≤ ε   λmin( H(θ+δθ)ᵀ W H(θ+δθ) )

This changes the answer, and I suspect it changes it in cross-hatch's favour — which leads to the more interesting point.

The loop-closure constraint pA(tc) = pB(tc) is written as a hard equality at a single intersection. Treat it instead as many soft revisits in a factor-graph formulation, SLAM-style, and notice what that buys: the discrepancy at each revisit is a direct measurement of accumulated tracking drift. The trajectory becomes its own calibration signal.

Sweep A Sweep B revisit same physical point, twice gap = measured drift more revisits = tighter drift bound

Each crossing is a free measurement of tracking error. Designed revisits turn the sweep into a self-calibrating instrument.

This is a second, independent axis of value that has nothing to do with sampling coherence — and it is arguably the stronger claim, because it attacks the risk your brief rates as most likely to be fatal.

Refinement 5

Two more borrows: blue noise, and cross-simulator validation

Blue-noise and low-discrepancy sampling, from computer graphics. Graphics spent thirty years on precisely this problem — placing samples that behave like randomness but never clump, with a designed power spectrum. A sum of K harmonics is a very impoverished sample spectrum: a few spikes. Low-discrepancy blue-noise construction gives a principled way to shape the spectral profile of the sampling while retaining the smoothness a hand requires. In ultrasound this machinery currently appears only in scatterer simulation, never in acquisition design.

Sum of K harmonics frequency →   a few spikes, everything else empty
Blue-noise profile low frequencies suppressed, no clumping, no gaps

Cross-simulator validation, as a guardrail on the optimiser itself. The pre-mortem risk is that the twin flatters trajectories designed inside it. Make that concrete: optimise Ω on twin A, then evaluate the resulting trajectory on twin B built with a different PSF model, speckle statistics and attenuation profile. If the advantage survives the swap it is real. If it grows as you idealise the simulator, you have caught the failure mode before it costs a year.

Three smaller adjustments

1 Make the one-sentence rule a formal budget rather than a principle — at most K ≤ 3 harmonics and two named motions — so it is enforced inside the optimiser rather than checked afterwards. Then test it at the phantom stage: can a sonographer actually reproduce trajectory Ω? That small human-factors sub-study would make the clinical claim much harder to dismiss.
2 Carry the not-missing-at-random point into the twin. Uniform frame dropout is optimistic. Real gaps cluster exactly where scanning is hard, which is where the data matters most — so model contiguous and structured dropout, not just random.
3 Keep the coherence number as a cheap surrogate, not the objective. Compute μ if it is convenient, but validate it against the empirical recovery curve. For deterministic, structured, continuous sampling of a non-sparse medium, no single scalar predicts recovery well — the curve is the evidence.

9 Augmenting the Instrument

The natural next thought is to bolt a second sensing modality onto the probe — infrared was the starting suggestion. The honest physics kills the obvious version and points at a much stronger one, and the real prize turns out not to be adding a better picture at all, but adding a channel that fixes the geometry.

First

Why plain infrared can't work — and what your instinct was actually reaching for

Thermal infrared is a surface measurement, killed twice over: water is essentially opaque at those wavelengths so nothing emitted from more than a fraction of a millimetre deep ever escapes, and the heat that does conduct to the skin has been through thermal diffusion — a spatial low-pass filter. A vessel 3 cm down produces, at best, a broad warm smudge with no depth information. Terahertz shares the problem; microwave penetrates but at centimetre resolution. All surface-or-blurry — rule them out.

But the instinct is right. What makes a vessel stand out under near-infrared light is optical absorption by haemoglobin. The trick is not to detect the light coming back — it's to detect the sound it makes. A nanosecond laser pulse is absorbed, the absorber expands, and that expansion emits an ultrasound pulse your existing transducer already hears. That is photoacoustic imaging: optical contrast at ultrasonic depth and resolution, on the same probe and the same RF chain. Your idea was sound; it just needs the signal to come out acoustically.

Second

The reframe: fuse for invariance, not for more pixels

Your brief's Risk #2 is that there is no single view-invariant volume — B-mode intensity depends on angle, pressure, gain and vendor processing, so you are fitting one volume to measurements that disagree by construction. The useful, unexploited fact: several channels from your existing probe are far more view-invariant than B-mode. Stiffness (elastography) is a tissue property that does not care about insonation angle. Flow (Doppler) depends on angle through a known, correctable cos θ law. Photoacoustic contrast is optical absorption, essentially independent of acoustic viewing angle.

So the fusion that helps is not "more data for a prettier picture." It is a view-invariant scaffold that pins down where things are, with view-dependent B-mode supplying the detail of what they look like.

Third

The borrow that makes it clean: cross-gradient joint inversion

Geophysicists solved exactly this decades ago. Fusing seismic, gravity and electromagnetic surveys, they never assume the modalities measure the same quantity — they don't. Instead they impose a structural coupling: the cross-gradient constraint, which requires the two property fields to have parallel gradients. In plain terms — the boundaries coincide, even though the values are unrelated. An organ edge is an edge in stiffness, in flow, in absorption and in echogenicity, but the numbers either side have nothing to do with each other.

B-mode (view-dependent) dense detail, wobbly edges values: echogenicity Stiffness / flow (view-invariant) sparse, but a crisp boundary values: unrelated to B-mode + Cross-gradient fused detail, anchored to a firm edge ∥∇x_Bmode × ∇x_stiff∥² → 0

The coupling term forces edges to align without ever assuming the two channels agree on values. Schematic.

Formally you add a term like ‖∇xB-mode × ∇xstiffness‖² to the same objective from section 8. It is well-behaved, it makes no false assumption that the modalities agree, and it lets a sparse view-invariant channel geometrically anchor a dense view-dependent one. As far as I can find it is unused in freehand 3D ultrasound — and it composes with everything already in this report as just one more term in the variational objective.

Fourth

Ranked by value per unit of pain — given solo + RF + phantoms

ModalityAttacksHardwareHonest read
Doppler / flowView-invariance, vessel identityNoneBest value on the list — same probe, same RF, data you already have. Angle dependence is known, so correctable.
ElastographyView-invariance, boundariesNone–lowA genuinely angle-independent property field. The ideal cross-gradient partner.
Force / pressure at faceDeformation (nonrigid risk)LowConverts your worst unmodelled confounder into a measured covariate — exactly the "pressure covariate" your brief asks for. Simulatable.
Depth / structured-light cameraSkin manifold, surface deformationLow–modFixes a simplification in your own generator: you assume the skin is flat (z = 0). It isn't. This measures the real surface 𝒮.
IMUPose, fast motionVery lowUseful but thoroughly worked. Not a contribution on its own.
Probe camera (VIO)Pose, driftModerateWorks well — and because it works well, it is crowded.
PhotoacousticVessels, oxygenation, invarianceHighThe correct form of the infrared idea, highest ceiling — but breaks the solo/phantom constraint and the freehand-PA lane is already active. Paper three.
Thermal IR / THz / microwave / EITRule out. Surface-only, or resolution measured in centimetres.
Fifth

The most novel angle: co-design the sweep and the sensing

Every fusion paper found fuses sensors for pose estimation — reconstruct better from whatever was collected. None asks the question this report is built around: what sweep maximises the information of the combined operator? The optimal-experimental-design machinery from refinement 1 does not care whether the rows of H come from one modality or four. So you can compute the D-optimal trajectory for a probe that also carries a force sensor and a Doppler channel — and the best sweep for a fused instrument is generally not the best sweep for B-mode alone. That question is unclaimed, it is pure simulation, and it extends your contribution rather than competing with it.

What I'd actually do

Add zero hardware first (Option F), then one cheap sensor.

Take Doppler or elastography as a second channel from the probe you already have, couple it to B-mode by cross-gradient joint inversion, and treat the angle-invariant channel as the geometric scaffold. It is simulatable in your twin, needs no calibration rig, does not break the solo constraint, attacks Risk #2 directly, and is one extra term in the objective. Then add the force transducer — deformation is the risk your brief rates most likely fatal, force is its direct cause, and it is the only sensor on the list you can model faithfully in simulation before buying it. Hold photoacoustics as the long game — your RF access already puts you closer to it than most, and it is the natural home for Option C's "reconstruct the medium, not the picture" ambition.

10 Watch Out For

The pre-mortem: assume in twelve months this failed. Here is why, with the sign that would have told you early.

Failure mode Pattern Why it bites Early-warning sign
Building the full forward model first Boiling the ocean Elevational PSF + scan conversion + RF likelihood + pose refinement is six months of correct engineering. It is a dependency of the claim, not the claim. Do it after the probes say it's worth it. It's week 3, you're still debugging the PSF model, and you have not produced a single recovery curve.
Simulation-flattered results Paving the cow path Your twin will reward the sampling pattern you designed, because the simulator's assumptions are your solver's assumptions. The result looks spectacular and does not survive a phantom. The designed trajectory's advantage grows as you make the simulator more idealised. That is diagnostic, and it's fatal if ignored.
A protocol no hand can perform Optimising the wrong stakeholder Jitter that breaks probe contact, or requires angular precision a sonographer can't hit, makes the whole acquisition-reduction claim fictional — and unpublishable in a clinical venue. You cannot describe the sweep to a sonographer in one sentence.
Uniform random frame dropout Optimistic assumption Real missing frames are missing not at random — they cluster exactly where scanning is hard, which is where you most need the data. Uniform dropout systematically overstates performance. Performance collapses the moment you switch from random dropout to contiguous or structured dropout.
Chasing PSNR / SSIM Optimising one metric Your own brief flags that these improve while geometry and diagnosis worsen. It is also the trap Probe 3 is designed to expose — don't fall into it while building the detector for it. Your headline figure is a side-by-side image comparison rather than a task metric or a recovery curve.
Scope drift into Option C Chasing the shiny option The RF scatterer-field idea is the most intellectually attractive thing in this report. It is also the one with 40% confidence. It will pull at you continuously. You're doing RF forward-modelling work that no probe asked for.

11 If This Were Reversed

The strongest case against this recommendation

Sampling design may be a lever ultrasound structurally cannot pull. MRI can jitter k-space because the sampling pattern is software. Seismic can jitter shots because a truck moves. But a freehand sweep is executed by a human hand on a deformable body in contact with the probe — the achievable "jitter" may be so constrained by contact, comfort and anatomy that the coherence improvement is negligible. If so, the field's focus on priors isn't a blind spot; it's a correct reading of where the degrees of freedom actually are.

And there is a publication-politics risk. Reviewers in medical imaging pattern-match to novel methods. A paper whose contribution is "a better sweep protocol and a borrowed validation metric" can read as incremental to someone skimming, however much stronger the science is. Option A produces an artifact that looks like what the venue expects.

This matters enough to bound rather than dismiss. It is the single assumption that, if wrong, invalidates the whole recommendation.

Risk-bounding checklist — what makes the move safe anyway

1 Put the hand-achievability constraint into the trajectory generator on day one, not as a later realism pass. If the designed sweeps must be physically performable from the start, the "ultrasound can't pull this lever" objection is answered by construction rather than argued about.
2 Set the dampen threshold before running Probe 1 and hold yourself to it. Under 5% separation inside the noise band means the objection was right — and you will have found that out in four weeks for the cost of some compute, not in eighteen months.
2b Run the unconstrained oracle first (refinement 2). It measures this exact objection directly: if the ceiling with no kinematic limits is barely above a smooth raster, then ultrasound genuinely cannot pull this lever, and you know it in week two rather than inferring it from a failed experiment later.
3 Run Probe 3 first. The trust instrument is valuable under every outcome, including the one where sampling design fails. It converts a possible null result into a publishable methodology contribution.
4 Frame the paper as a method plus protocol, not a protocol alone. Attach the sampling result to the physics-aware reconstruction from Option A — which you build anyway — so the artifact reads as the novel method reviewers expect, with the sampling contribution as its foundation.
5 Keep Option C alive as a stated follow-on in the discussion section. It costs nothing, stakes the RF territory publicly, and preserves the moat for paper two.
6 Get to a physical phantom by month three. Everything above is simulation, and simulation is where this recommendation is most likely to fool you.