Current artificial intelligence evaluation often relies on static benchmarks and task-specific performance metrics, which are useful for system comparison but limited in assessing developmental progression, cognitive architecture, adaptive transfer, and real-world applicability. Building on developmental approaches to AGI testing, including our earlier conceptual proposal, this article develops an operational formulation of the General–Specialized–Applicable (GSA) framework for artificial general intelligence (AGI) evaluation. The framework organizes evidence into three non-interchangeable stages. The General stage, which assesses foundational capacities for open-ended generalization, value-oriented regulation, and autonomy; the Specialized stage, which evaluates stable domain-specific competence; and the Applicable stage, which examines whether such competence can be deployed safely and robustly in realistic environments. The article further introduces a dynamic task-generation pipeline and a provisional operational rubric for interpreting stage-specific evidence. An illustrative case study using household tasks in a simulated embodied environment illustrates how GSA can provide diagnostic information beyond aggregate benchmark scores by identifying capability gaps in current multimodal large language model (MLLM) agents. Rather than offering a final universal standard, the GSA framework provides an evaluation-oriented structure for connecting benchmark performance with architectural readiness, specialization, and deployment-level applicability.