Most of these ICs have a lot of FLASH/RAM that contribute to the die size, the FPU-D is perhaps 3x size of FPU-S, I couldn't get the ARM guys to share specific info.
One thing the OP didn't make clear was whether it is acceptable to do the floating-point in software.
All the compilers provide software double-precision floating-point where it's needed. It will increase execution-time (hence power-consumption), and use up more FLASH. But often this is a good overall solution.
But even the smallest stm32l0 can do it. A quick check found stm32l052T8 is 2.61 x 2.88 mm but I don't know if that's the smallest.
Yes, we need to shutdown quick, so hardware is the best answer. In the end, I Had to compromise, with the L433-48pin, will have to calculate the double from the Single Floating Point Hardware.