Actually a single camera is all you need. I think it’s fair to say that the only thing stereo gets you is scale. But both cameras and lidar have their place in sensing systems, and getting more experience with either is useful.
If you’re interested in reconstruction from images check out Meshroom and Nerf Studio
Scale is the one thing stereo doesn't get you compared with sequential mono images, unless you have some fancy lens model that lets you derive scale from nonlinearities in the lens. Is that something we do now? I always wanted to try out monocular SLAM with a fisheye lens.
With 2 mono images you can figure out that an object is twice as big as an other, but you can't tell the size of any objects (= you don't know the scale).
With a stereo image you know the distance between the lenses, which allows you to know the size of the objects (= you know the scale).