Methodology
The STEP methodology: how we assess, teach, and measure professional communication.
STEP has been built inside The Hindu Group, which has published in English in India since 1878. This page sets out our methodology in full — the pedagogical tradition we work in, how our assessments are designed, how a lesson and a coach call are structured, and how we measure whether any of it worked.
One standard governs everything below: if we cannot defend a claim in print, we do not make it.
The method at a glance
Five parts. One connected method.
Align
Base aims, lessons and outcomes on communicative teaching principles and CEFR standards.
Assess
Establish the starting point through calibrated, adaptive assessment.
Teach
Build language through interaction, instruction, practice and feedback.
Reinforce
Connect app practice, coaching and classroom learning in one loop.
Measure
Reassess progress and report outcomes against the same external standard.
Applied to the learner
The same method, applied to different needs.
Individuals
A personal diagnostic shapes practice, coach calls and progress checks.
- Diagnostic starting point
- Personalised app practice
- Specific coach feedback
- Progress measurement
Organisations
A cohort baseline informs level-based delivery and defensible outcome reporting.
- Cohort baseline
- Level-appropriate delivery
- Blended reinforcement
- Outcome reporting
Built through practice
Look under the bonnet.
The principles, assessment design, teaching structure and evidence behind STEP — organised so you can go directly to the part that matters to you.
Our pedagogical foundation: communicative language teaching
STEP’s teaching is grounded in communicative language teaching (CLT) — the approach used by leading language programmes worldwide. CLT’s central claim, borne out by decades of second-language acquisition research, is that language is acquired through meaningful use, not through memorising rules about it. A learner who practises structuring a client update acquires the past tense faster — and retains it longer — than one who studies the past tense as an abstraction.
Three CLT principles run through every STEP course, lesson, and coach interaction:
Contextual. Language elements are introduced the way they appear in real life — the past tense through “what I did last weekend,” a formal register through an actual email. Anchoring new language in a learner’s existing experience of the world (their schemata) makes acquisition measurably more effective, keeps relevance visible, and builds the positive attitude towards learning that sustains a programme.
Skill-based. We teach the four communicative skills — listening, speaking, reading, writing (LSRW) — explicitly and separately, because Indian secondary and tertiary education largely does not. Each lesson targets one core skill, supported where needed by grammar and vocabulary.
Level-based. Language acquisition is most effective when input is pitched at the right level; material that is too hard or too easy demotivates equally. Every learner is assessed before instruction begins, and the programme they receive is pitched to their measured level.
Sub-skills and functional language
Beneath the four skills, our method works at the level of sub-skills. For the receptive skills — listening and reading — improvement comes from deliberately building components such as listening for gist, listening for specific information, reading for inference, and reading for opinion. Each lesson targets two or more sub-skills; as the building blocks develop, the skill develops.
For the productive skills — speaking and writing — we teach functional language: the ready-to-use phrases and structures needed for a specific communicative job. Asking for information. Clarifying. Seeking permission. Persuading. Registering a complaint. Rather than the nuts and bolts of grammar in the abstract, the learner is equipped with components that have an immediate effect on real communication — and the grammar is acquired through their use.
Cognitive design
Lessons are also deliberately sequenced against Bloom’s taxonomy of cognitive processes. Grammar lessons work at the remember and understand levels, surfacing the substantial dormant knowledge most Indian learners already hold. Every lesson teaches target language in a real-life situation, so application is built in. Error-spotting and good-example/bad-example work triggers analysis and evaluation. And every productive-skills lesson ends with the learner creating language — speaking or writing something of their own. The taxonomy is not decoration; it is why lessons end in production rather than in a quiz.
Scaffolded progression
Skills are sequenced so that foundations precede demands. In writing, for example, sentence control, coherence, and organisation are established before learners attempt advanced genres — introducing report-writing to a learner without sentence control places disproportionate cognitive load on them and teaches mostly frustration. This scaffolding principle, applied across all four skills and across CEFR levels, is the same one we applied when the Tamil Nadu State Council for Higher Education engaged us to help redesign the state’s undergraduate Part II English curriculum.

Curriculum design in practice
The three CLT principles
Contextual
Language elements are introduced the way they appear in real life, anchored in the learner’s existing experience of the world.
The past tense through “what I did last weekend.” A formal register through an actual email.
Skill-based
Listening, speaking, reading and writing are taught explicitly and separately, because Indian secondary and tertiary education largely does not.
Each lesson targets one core skill, supported where needed by grammar and vocabulary.
Level-based
Input is pitched at the right level, because material that is too hard or too easy demotivates equally.
Every learner is assessed before instruction begins, and the programme is pitched to their measured level.
The CEFR backbone
Everything at STEP — course aims, lesson objectives, assessment items, score reports, certificates — is aligned to the Common European Framework of Reference for Languages (CEFR), the international standard for describing language ability. The product of over twenty years of research, the CEFR describes six proficiency levels, A1 to C2, each defined by “can-do” statements: performance-related descriptions of what a learner at that level can actually do, developed by the Association of Language Testers in Europe (ALTE).
What a learner at each level can do (abridged)
Understands demanding, longer texts and implicit meaning. Expresses ideas fluently and spontaneously; uses language flexibly for social, academic and professional purposes.
Understands complex text including technical discussion in their field. Interacts with fluency that puts no strain on either party. Produces clear, detailed text and argues a position.
Handles most situations at work, study, and travel. Produces simple connected text on familiar matters; describes experiences and gives reasons for opinions and plans.
Communicates in simple, routine tasks on familiar matters. Describes their background and immediate environment in simple terms.
Understands and uses familiar everyday expressions and basic phrases. Interacts simply if the other person speaks slowly and is prepared to help.
The learning aims of every STEP skill lesson are derived directly from the can-do statements at the relevant level. For ease of use, STEP subdivides the CEFR bands into a 12-level scale — finer gradations that let a learner (and an employer or placement cell) see movement within a band, not just across bands. The CEFR also specifies the contexts in which language is used — occupational, public, educational — and our items and lessons are written against those contexts, which is why so much of our material is set in offices, interviews, service settings, and workplaces rather than in abstract exercises.
Why insist on an external standard? Because self-referential scores are unfalsifiable. A STEP level means the same thing to a recruiter in Chennai, a placement cell in Coimbatore, and a global capability centre calibrating a hiring bar — and it can be checked against a published international framework rather than taken on our word.
Assessment design
Assessment is where STEP started and remains the anchor of the platform. We conduct over 200,000 assessments a year, and every engagement — enterprise, institutional, or individual — begins with one. Four design principles govern how STEP assessments are built.
Principle 1 — Global standards, externally referenced
Every item, level boundary, and score report maps to the CEFR as described above. We do not grade against a proprietary scale that only we can interpret.
Principle 2 — Calibrated items testing specific skills and sub-skills
Each test item is written to assess a specific skill, at a specific proficiency level, in a specific context. A B2 spoken-production item, for instance, takes the CEFR descriptor — “can deliver announcements on most general topics with clarity, fluency and spontaneity” — and situates it: you work for the railways; a train has been cancelled; prepare the passenger announcement, three minutes to prepare, three minutes to record. The skill, the level, and the context are all deliberate choices, not accidents of the question writer’s imagination.
Items are drawn from a large, actively managed bank: our total pool exceeds 14,000 items, of which roughly 3,100 calibrated items are live in assessments at any time, with 700–800 new items added and calibrated each year based on how items perform with real test-takers. Reading sections, for example, move through five task types — from filling gaps in real correspondence to answering questions on persuasive and argumentative articles — testing sub-skills from reading for gist through reading for inference.
Principle 3 — Adaptive testing
STEP assessments are computer-adaptive: the difficulty of each question adjusts to the test-taker’s performance on the previous ones. This is how a proficiency level can be fixed with high confidence in 15–20 questions per section rather than 60 — the test converges on the learner’s level instead of marching every learner through the same paper. Adaptivity requires every item to be objectively auto-gradeable, which is fully achievable for reading and listening and largely achievable for speaking and writing, with the remainder handled by the open-ended tasks below.
Principle 4 — AI-based evaluation of open-ended speech and writing
The hardest part of language assessment at scale is grading what a learner produces freely. STEP handles this in two stages. In the adaptive portion, structured speaking tasks are evaluated at phoneme level for pronunciation and fluency against expected responses. Then, in the open-ended portion, the learner speaks for 1–2 minutes and writes 100–150 words in response to prompts selected at their tentative level — and the writing is evaluated on four dimensions: ability to stay on topic, readability, grammatical accuracy (error rate against total word count), and lexical range. The open-ended score is combined with the adaptive score to produce the final result. Speech evaluation uses specialist third-party engines alongside our own frameworks; the design principle is that no single automated measure decides a score alone.
Test formats
STEP Plus is the flagship assessment: a comprehensive, proctored evaluation of 60 minutes across 23 tasks, with 15 scored responses in each of listening, speaking, reading, and writing. It is the instrument behind STEP certification and formal benchmarking — the score an employer calibrates a hiring bar against and a learner carries into the job market.
Everything else is a configuration, not a separate product. From the same calibrated item bank and the same CEFR output standard, we build customised assessments to fit the job at hand — a rapid screening pass for high-volume hiring, a short diagnostic to baseline a cohort and place learners into level-appropriate instruction, or a client-specific instrument weighted towards the skills and contexts that matter for a particular role. Duration, task mix, and skill weighting change; the standard behind the score does not. This is deliberate: a screening result, a diagnostic profile, and a certification are comparable because they come off one infrastructure — not three tests that happen to share a brand.
| 200,000+ assessments a year | 14,000+ item pool · ~3,100 live | 700–800 new calibrated items each year | 15–20 questions to fix a level, adaptively |
|---|
Score report · overall score, four skill bands, CEFR level and skills breakdown
STEP Plus certificate · the same CEFR-mapped scale, issued against a certificate number
How instruction is structured
Assessment tells us where a learner is; instruction is how they move. STEP delivers instruction three ways — self-paced online lessons, classroom training, and live 1-on-1 coach calls — and combines them in blended programmes. Each is a distinct discipline with its own structure, and each is set out below.
The online lesson
Every productive-skills online lesson follows a four-part arc: interaction → instruction → practice → feedback. The target language is first met inside a natural interaction — a conversation, an email, an announcement — never as a rule on a slide. The instructor then draws attention to the phrases, structure, and tone that did the work in that interaction, and extends them. The learner practises. The learner gets specific feedback on the practice, including why wrong answers were wrong — because the explanation of an error is where the learning happens.
The anatomy of an online lesson
Interaction
The target language is first met inside a natural interaction — a conversation, an email, an announcement — never as a rule on a slide.
Instruction
The instructor draws attention to the phrases, structure and tone that did the work in that interaction, and extends them.
Practice
The learner practises the target language.
Feedback
Specific feedback on the practice, including why wrong answers were wrong — because the explanation of an error is where the learning happens.
Tuned per skill
Speaking
Opens with a filmed interaction; closes with the learner producing the function themselves.
Writing
Opens with a model text — what good looks like — before deconstructing its structure, tone and phrasing.
Listening
Sets context first, uses audio from natural situations, and questions across sub-skills.
Reading
Pairs authentic passages with sub-skill questions and feedback that builds reading strategies, not just answers.
The arc is tuned per skill. Speaking lessons open with a filmed interaction and close with the learner producing the function themselves. Writing lessons open with a model text — what good looks like — before deconstructing its structure, tone, and phrasing. Listening lessons set context first, use audio from natural situations (announcements, tours, radio), and question across sub-skills. Reading lessons pair authentic passages with sub-skill questions and detailed feedback that builds reading strategies, not just answers. Grammar lessons cover the core competencies with particular attention to the patterns Indian learners most often stumble on — subject–verb agreement, prepositions, articles — surfaced through interaction rather than drilled through rules.
The classroom
Classroom training is where STEP delivers at cohort scale, and it is a distinct discipline from coaching — not a coach call with thirty people in the room. Three design decisions define it.
First, no mixed-ability teaching. Every classroom programme opens with a proctored baseline assessment, and the cohort is grouped into level-based batches — beginner, intermediate, advanced — before the first session. Teaching to a mixed-ability room fails everyone in it: the material is simultaneously too hard for the bottom third and too easy for the top. Batching by measured level is the classroom expression of the same level-based principle that governs our self-paced courses.
Second, activity-based and communicative throughout. A STEP classroom session is built on the CLT progression from controlled practice to freer communicative tasks. The trainer introduces the target function in context and runs controlled practice — structured, scaffolded, high-correction. As competence builds, the scaffolding comes away and the session moves to freer production: role plays, simulated group discussions run against real evaluation rubrics, mock interviews (HR and technical, 1-on-1), presentations to the room. Learners spend the majority of session time producing language, not receiving it. Units are integrated — grammar, reading, listening, speaking, and writing reinforce each other within a unit rather than running as parallel silos — and programmes build to a course-end showcase where learners perform what they can now do, assessed against the same rubrics used throughout.
Third, the classroom never works alone. Contact hours are paired with an individually-assigned online layer — each learner’s self-paced pathway is generated from their own diagnostic, so classroom time reinforces the cohort-level curriculum while the online hours attack each individual’s specific gaps. Printed workbooks extend the classroom beyond contact hours: self-explanatory by design, restating key concepts and providing practice activities a learner can complete with minimal teacher intervention.
Who stands in front of the room matters as much as the method. STEP classroom trainers are CELTA-certified language teaching professionals and industry veterans — people who have run corporate recruitment and client communication themselves — because placement and workplace training requires both credibility about the destination and craft in the teaching.
The coach call
Where the classroom moves a cohort, a coach call moves one person. A live 1-on-1 call applies the same lesson arc to a single learner’s specific gaps: the coach works from the learner’s assessment profile, the record of previous calls, and above all the assignment the learner completed since the last call. The call is never generic. A typical call runs 30 minutes and moves through the four stages below; calls on the IELTS Speaking course and in corporate programmes run 45 minutes, where a wider scope of work needs the extra time. A coach follows the learner, so a call that needs longer on practice takes it from somewhere else:
- Feedback on the assignment. Every call ends with an assignment; every next call opens with feedback on it. The coach names what the learner got right and where the errors were — specifically, against the language the learner actually produced. This loop — an assignment set, then feedback on that assignment — is the single most important mechanism in the coaching model.
- Mini-teach, where the assignment calls for it. Where the assignment exposes a particular area of weakness — a tense, a register, a pronunciation feature — the coach teaches to it there and then. The stage is conditional: it runs when the learner’s own work has shown something to fix, not on a schedule.
- The input session. The coach introduces the topic for that call and the communicative function it turns on — disagreeing with a senior, opening a client call, structuring bad news — inside a realistic scenario drawn from the learner’s own work or study context. Practice sits inside this stage and moves in one direction. Guided practice comes first: the learner produces the language with the coach scaffolding — prompts, corrections at the point of error, second attempts. Freer practice follows: the scaffolding comes away and the learner runs the scenario end-to-end with minimal or no support — a mock announcement, a role-played conversation, a summarised argument — while the coach observes against the four graded signals. What changes between the two is how much support is in the room.
- The next assignment. The coach sets a specific assignment shaped by what the call revealed, and assigns the app modules that practise exactly that. The learner does the work before the next call — and the next call opens with feedback on it.
The structure of a coach call
Feedback on the assignment
Every call ends with an assignment; every next call opens with feedback on it — what the learner got right, and where the errors were.
Mini-teach, if the assignment calls for it
Where the assignment exposes a particular weakness, the coach teaches to it there and then. Conditional, not scheduled.
The input session
The call’s topic and target function, inside a realistic scenario. Practice runs here: guided first, with the coach scaffolding, then freer, with minimal or no support.
The next assignment
A specific assignment shaped by what the call revealed, plus the app modules that practise exactly that.
A typical call30 min

The blended loop
In blended programmes the modalities are not parallel tracks; they form a loop. The classroom and the coach build skill. AI-powered assessment surfaces each individual’s gaps. Personalised online practice reinforces exactly those gaps between sessions. Endline reassessment confirms — or denies — the uplift. Each element covers what the others cannot: a classroom cannot give thirty learners individual pronunciation feedback; an app cannot notice that a learner’s hesitation comes from confidence rather than vocabulary. The loop can.
The blended loop
The classroom and the coach build skill
Contact hours and live 1-on-1 calls move the cohort and the individual.
Assessment surfaces the gaps
AI-powered assessment surfaces each individual’s gaps.
Online practice reinforces them
Personalised online practice reinforces exactly those gaps between sessions.
Reassessment confirms the uplift
Endline reassessment confirms — or denies — the uplift, on the same CEFR-mapped scale.
And the loop runs again
Each element covers what the others cannot. A classroom cannot give thirty learners individual pronunciation feedback; an app cannot notice that a learner’s hesitation comes from confidence rather than vocabulary. The loop can.
What we measure, and what we report
Four signals are graded every cohort, with audio evidence retained for review:
- Fluency — words per minute, hesitation, recovery after a stumble.
- Grammar — tense control, register, agreement, and the patterns Indian speakers most often stumble on.
- Vocabulary — range and precision in workplace, exam, or test-day contexts.
- Pronunciation — sound by sound, at phoneme level, with retained audio.
Measurement is pre/post by design: a proctored baseline before instruction begins, an endline after it ends, both on the same CEFR-mapped scale. For organisations, this becomes the outcome report — CEFR movement per learner and per cohort, engagement and completion, sub-skill breakdowns, and, where the data is shareable, placement readiness or job-progression deltas. Institutions additionally get a live dashboard with cohort- and individual-level drill-down. The reporting exists for one reason: the person who bought the programme should be able to defend the spend with data, and the learner should be able to defend the credential.

Cohort and learner reporting · client dashboard, identities removed
The evidence that it works
Methodology claims are cheap; measured outcomes are not. Three kinds of external validation matter to us.
Measured cohort outcomes. In a recent academic-year engagement with a leading South Indian engineering institution — a multi-year partnership now in its fourth year — 1,567 first-year students across 14 branches completed our 50-hour blended programme; 1,527 of them moved up at least one STEP proficiency level between baseline and endline, with 81% completing the online component. Nearly every learner who finished showed measurable uplift on the same externally-referenced scale they started on.
Government and academic selection. STEP’s learning programme was selected under the AICTE/MHRD National Educational Alliance for Technology (NEAT) programme — a selection made specifically on the basis of measurable learner impact and the use of AI tools to drive improvement. Separately, the Tamil Nadu State Council for Higher Education engaged STEP’s Content and Pedagogy team in the redesign of the state’s undergraduate English curriculum — CEFR-aligned, skills-based, and performance-assessed on precisely the principles described on this page.
Institutional and enterprise adoption at scale. Organisations that have used STEP for assessment and training include TCS, Cognizant, L&T, Bank of America, Ford, and ICICI Prudential, and institutions including ISB, Christ University, Manipal Academy, and the Shiv Nadar Foundation — buyers with the sophistication to test claims before renewing.

Measured cohort movement · diagnostic to certification by skill, client dashboard with identities removed
From the Content and Pedagogy team
English is accessible to everyone. Everyone can learn to speak with confidence.
We have taught in engineering colleges and on contact-centre floors. We have coached vice-presidents and first-year support reps. We have watched students who failed a Class X English board go on, three years later, to train the new joiners.
Everything above — the CLT foundation, the CEFR backbone, the adaptive assessment, the lesson arc, the coach loop — exists in service of three plain observations. People do not need to be told they are smart; they need to be told, specifically, what went wrong in the work they just did. They need someone who will look at the next piece of work and tell them whether it went better. And they need to be taken seriously. The platform sets the practice. The coach reads the work and says what changed. And we treat every learner — engineer, exam aspirant, support rep, study-abroad hopeful — as if the work they are doing is real, because it is.
— STEP CONTENT AND PEDAGOGY TEAM, CHENNAI