A document-processing pipeline has two stages: (1) high-volume extraction of simple fields from millions of short documents, and (2) low-volume synthesis of complex cross-document analysis reports. Which model assignment best balances cost, latency, and capability?