You record the video from a matrix of cameras each at a slightly different pov, then use a light-field algorithm to generate novel images. As you said you only solve occlusion if your cameras are separated enough but in my case I only expect/want the user to twitch his head around. We can add more cameras if this tech starts get adopted.