QFI to Wind Down Operations.

Read the Announcement

Proficiency vs. AI: Navigating the New Era of Language Assessment

Jul 16, 2026

By Sarab A.

Proficiency-based assessments evaluate what language learners can spontaneously do with language in real-world, unrehearsed situations, rather than measuring passive knowledge of grammar or vocabulary lists.

Grounded in major frameworks like the American Council on the Teaching of Foreign Languages (ACTFL) Guidelines and the Common European Framework of Reference for Languages (CEFR), these assessments focus on three primary modes of communication: Interpersonal, Interpretive, and Presentational. Rather than testing classroom-specific performance, a true proficiency assessment serves as an independent, curriculum-free evaluation of a learner's functional ability. This approach views the learner as an active "social agent" who dynamically integrates grammatical, sociolinguistic, and strategic competencies to negotiate meaning, complete tasks, and achieve non-linguistic outcomes. Ultimately, the defining characteristic of proficiency assessment is its consistency with target language use (TLU)—meaning evaluation tasks must closely mirror the authentic, real-world scenarios learners navigate outside the classroom (Bachman & Palmer, 1996; Canale & Swain, 1980; Council of Europe, 2001; Swender et al., 2012).

When designing effective proficiency-based assessments, educators rely on three core principles. First, the evaluation must be strictly criterion-referenced, meaning student performance is measured against fixed, global benchmarks—such as the ACTFL Novice, Intermediate, and Advanced levels—rather than being graded on a curve (Clifford, 2016). Second, the assessment must target spontaneous language control in unrehearsed situations, independent of specific classroom curricula (Swender et al., 2012). Finally, the primary metric of success is functional capability: whether the student effectively communicated their message and completed the communicative task at hand (Council of Europe, 2001).

To illustrate these principles, authentic language tasks can be aligned across different communicative modes and proficiency levels:

  • Novice Interpersonal: In a task like The New Roommate, students interview each other spontaneously to find a compatible match for a summer program by asking and answering simple, discrete questions about daily routines.
  • Intermediate Interpretive: Students listen to an authentic weather forecast to determine what to pack for a trip, requiring them to extract highly structured, paragraph-length factual information like temperature shifts and weather warnings.
  • Advanced Presentational: After arriving abroad to find their car rental unavailable due to an agency glitch, students record a formal, structured voicemail for corporate headquarters that seamlessly weaves together a past narration of their booking, a present description of their logistical gridlock, and a future contingency plan.

While proficiency-based assessments have been successfully utilized within modern language pedagogy since the 1980s, the sudden emergence of generative artificial intelligence introduces a pivotal turning point for the field. As AI integrates into education, widespread concerns have emerged regarding how a high student reliance on automated tools might negatively impact independent language acquisition. This shift brings educators face-to-face with one critical question: do proficiency-based assessments now play a more vital role than ever before in validating true, unassisted human capability, and is AI ultimately a supportive friend or a formidable foe in this evaluative process?

Spontaneously completing a real-world task in an unrehearsed situation sounds like the ultimate defense against academic AI reliance, assuring language educators that students are producing authentic, unassisted work. Yet, a critical step back reveals that while this security is valid, the true challenge lies within the sophisticated evaluation phase. To truly measure proficiency, evaluating a task simply by whether it was successfully completed is insufficient; completion proves a goal was met, but fails to reveal how it was achieved or the specific proficiency level demonstrated. According to major frameworks like ACTFL and the CEFR, performance must be evaluated against a multi-dimensional matrix of criteria—such as global functions, text type, context, and comprehensibility—that thoroughly assess the quality, independent control, and sustainability of the student’s language. This remains a highly demanding process, even for seasoned educators.

Recent literature confirms that AI can significantly alleviate this evaluation burden while simultaneously optimizing the student experience. In a longitudinal study, Zhan and Yan (2026) found that integrating AI-supported feedback significantly enhanced language learners' overall "feedback literacy"—specifically improving their capacity to appreciate evaluation, manage affective emotional responses, and actively take action during revisions. Similarly, Kamelabad et al. (2026) noted that immediate, automated corrections drastically enhanced the user experience, lessened language anxiety, and fostered highly positive student perceptions of conversational interaction. At the same time, however, Kinder et al. (2026) caution that AI cannot independently replace human grading; language teachers frequently must modify automated outputs to inject personalized student trajectories and maintain trust.

This precise balance between automated feedback efficiency and essential human oversight is exactly what I observed over the past year when I integrated Speakology AI into my Intermediate Arabic Heritage classes. This immersive conversational platform, where students interact with AI-driven avatars in task-based simulations, offers an architectural framework that allows educators to design tasks aligned precisely with specific ACTFL or CEFR proficiency levels and sublevels, ranging from Novice to Distinguished. Beyond basic task completion, the tool provides a detailed breakdown of each interaction across a multi-dimensional matrix of criteria: Functions, Contexts & Content, Text/Discourse Type, Language Control (Accuracy), Vocabulary Use, Communication Strategies, and Cultural/Register Awareness. For each dimension, the system assesses student performance on a scale of meeting, below, or exceeding expectations, validating these ratings with concrete performance descriptors (e.g., identifying text discourse as "using strings of sentences with limited cohesion and connectors"). Upon concluding an interaction, the AI immediately delivers targeted, criteria-aligned formative feedback to the learner, offering actionable insights such as praising the student's capacity to "sustain a conversation and express ideas about future plans," while simultaneously providing precise directives to "work on using correct verb forms and gender agreements" and "increase vocabulary range to express ideas more precisely and avoid relying on code-switching."

The true pedagogical power of Speakology AI lies in its domain-specific architecture; unlike generalized AI platforms, it is designed exclusively for language pedagogy and learning. The "secret ingredient" of this platform is the agency it returns to the instructor during the task-design phase. By controlling the simulation parameters, the educator can intentionally embed the exact linguistic triggers required for a rigorous, level-appropriate proficiency assessment. This customized engineering allows instructors to capitalize on a powerful "three-in-one" AI workflow: achieving precise task alignment, facilitating immersive student interactions with an AI avatar, and instantly obtaining a detailed performance breakdown alongside targeted, criteria-based feedback. Furthermore, the platform's creators reinforce the vital role of the human instructor by allowing educators to append personal notes and manually evaluate the performance, ensuring that teacher expertise remains central to the assessment process.

Ultimately, the intersection of AI and language pedagogy does not diminish the value of proficiency-based assessment; rather, it elevates it. When intentionally leveraged through domain-specific platforms like Speakology AI, artificial intelligence transforms from a threat to academic integrity into a powerful pedagogical ally. By streamlining the labor-intensive demands of multi-dimensional evaluation and offering immediate, formative feedback, AI enhances the teacher's capabilities without replacing their essential expertise. In this evolving landscape, the interaction between human instructional design and AI-driven simulation ensures that proficiency assessments remain robust, authentic, and uniquely equipped to validate true human capability in an automated world.

Disclaimer: The views and opinions expressed in this blog are those of the author and do not necessarily reflect the views of Qatar Foundation International (QFI). While QFI reviews guest contributions for clarity and to ensure the content is valuable for our audience, the accuracy and completeness of the information are the responsibility of the author.

Sarab A.

Sarab A. is a Senior Lector II of Arabic and Director of the Arabic Language Program at Yale University.

References

Bachman, L. F., & Palmer, A. S. (1996). Language testing in practice: Designing and developing useful language tests. Oxford University Press.

Canale, M., & Swain, M. (1980). Theoretical bases of communicative approaches to second language teaching and testing. Applied Linguistics, 1(1), 1–47.

Clifford, R. (2016). A rationale for criterion-referenced proficiency testing. Foreign Language Annals, 49(2), 224–234.

Council of Europe. (2001). Common European Framework of Reference for Languages: Learning, teaching, assessment. Cambridge University Press.

Kamelabad, A. M., Turano, B., Lundin, M., & Skantze, G. (2026). Personalized language learning with an LLM chatbot: Effects of immediate vs. delayed corrective feedback. Frontiers in Education, 11, Article 1703664. https://doi.org/10.3389/feduc.2026.1703664

Kinder, A., Lee, J., & Moore, M. (2026). Enhancing learner-centered feedback with AI: teachers' practices and perceptions. Assessment & Evaluation in Higher Education, 51(3), 312–326. https://doi.org/10.1080/02602938.2026.2638920

Swender, E., Martin, C. L., Rivera-Martinez, M., & Tseng, R. (2012). ACTFL Proficiency Guidelines 2012. American Council on the Teaching of Foreign Languages.

Zhan, Y., & Yan, Z. (2026). Exploring the impact of a GenAI-supported feedback practice on student feedback literacy in L2 writing. Computer Assisted Language Learning. Advance online publication. https://doi.org/10.1080/09588221.2025.2605541

Loading...