A few Superbowls ago Intel demoed a supercomputer which could take in live video streams from hundreds of cameras, run a batch job for about 30 seconds, and then be able to synthesize video from arbitrary viewpoints above the crowd.
I think it worked by making a point cloud that fits the camera observations.
Obstruction is not a big issue in that use case, but if there was obstruction, the system could choose to ignore pixels that were obstructed when constructing the image.
Right. And with the initial cameras, which were neat but not breathtaking in their capabilities, you really only got a small ability to move around the scene. A larger sensor or a pair of these would offer greater capabilities for this application. I believe using an array of sensors was how their video camera worked, but I'm not 100%.
(I spent a lot of time reading about light field stuffs back when Lytro was first announced, I was also tangentially connected to a synthetic aperture radar project so it piqued a lot of my curiosity at the time. I haven't kept up with the state-of-the-art since then, though.)