Methodology

The STEP methodology: how we assess, teach, and measure professional communication.

STEP has been built inside The Hindu Group, which has published in English in India since 1878. This page sets out our methodology in full — the pedagogical tradition we work in, how our assessments are designed, how a lesson and a coach call are structured, and how we measure whether any of it worked.

One standard governs everything below: if we cannot defend a claim in print, we do not make it.

Built through practice

Look under the bonnet.

The principles, assessment design, teaching structure and evidence behind STEP — organised so you can go directly to the part that matters to you.

01 · ALIGN

Our pedagogical foundation: communicative language teaching

STEP’s teaching is grounded in communicative language teaching (CLT) — the approach used by leading language programmes worldwide. CLT’s central claim, borne out by decades of second-language acquisition research, is that language is acquired through meaningful use, not through memorising rules about it. A learner who practises structuring a client update acquires the past tense faster — and retains it longer — than one who studies the past tense as an abstraction.

Three CLT principles run through every STEP course, lesson, and coach interaction:

Contextual. Language elements are introduced the way they appear in real life — the past tense through “what I did last weekend,” a formal register through an actual email. Anchoring new language in a learner’s existing experience of the world (their schemata) makes acquisition measurably more effective, keeps relevance visible, and builds the positive attitude towards learning that sustains a programme.

Skill-based. We teach the four communicative skills — listening, speaking, reading, writing (LSRW) — explicitly and separately, because Indian secondary and tertiary education largely does not. Each lesson targets one core skill, supported where needed by grammar and vocabulary.

Level-based. Language acquisition is most effective when input is pitched at the right level; material that is too hard or too easy demotivates equally. Every learner is assessed before instruction begins, and the programme they receive is pitched to their measured level.

Sub-skills and functional language

Beneath the four skills, our method works at the level of sub-skills. For the receptive skills — listening and reading — improvement comes from deliberately building components such as listening for gist, listening for specific information, reading for inference, and reading for opinion. Each lesson targets two or more sub-skills; as the building blocks develop, the skill develops.

For the productive skills — speaking and writing — we teach functional language: the ready-to-use phrases and structures needed for a specific communicative job. Asking for information. Clarifying. Seeking permission. Persuading. Registering a complaint. Rather than the nuts and bolts of grammar in the abstract, the learner is equipped with components that have an immediate effect on real communication — and the grammar is acquired through their use.

Cognitive design

Lessons are also deliberately sequenced against Bloom’s taxonomy of cognitive processes. Grammar lessons work at the remember and understand levels, surfacing the substantial dormant knowledge most Indian learners already hold. Every lesson teaches target language in a real-life situation, so application is built in. Error-spotting and good-example/bad-example work triggers analysis and evaluation. And every productive-skills lesson ends with the learner creating language — speaking or writing something of their own. The taxonomy is not decoration; it is why lessons end in production rather than in a quiz.

Scaffolded progression

Skills are sequenced so that foundations precede demands. In writing, for example, sentence control, coherence, and organisation are established before learners attempt advanced genres — introducing report-writing to a learner without sentence control places disproportionate cognitive load on them and teaches mostly frustration. This scaffolding principle, applied across all four skills and across CEFR levels, is the same one we applied when the Tamil Nadu State Council for Higher Education engaged us to help redesign the state’s undergraduate Part II English curriculum.

Three colleagues in discussion at a whiteboard during a working session

Curriculum design in practice

The three CLT principles

01

Contextual

Language elements are introduced the way they appear in real life, anchored in the learner’s existing experience of the world.

The past tense through “what I did last weekend.” A formal register through an actual email.

02

Skill-based

Listening, speaking, reading and writing are taught explicitly and separately, because Indian secondary and tertiary education largely does not.

Each lesson targets one core skill, supported where needed by grammar and vocabulary.

03

Level-based

Input is pitched at the right level, because material that is too hard or too easy demotivates equally.

Every learner is assessed before instruction begins, and the programme is pitched to their measured level.

The three communicative language teaching principles behind every STEP course. Contextual: language elements are introduced the way they appear in real life, such as the past tense through “what I did last weekend” or a formal register through an actual email. Skill-based: listening, speaking, reading and writing are taught explicitly and separately, each lesson targeting one core skill. Level-based: input is pitched to the learner’s measured level, established by an assessment before instruction begins.
02

The CEFR backbone

Everything at STEP — course aims, lesson objectives, assessment items, score reports, certificates — is aligned to the Common European Framework of Reference for Languages (CEFR), the international standard for describing language ability. The product of over twenty years of research, the CEFR describes six proficiency levels, A1 to C2, each defined by “can-do” statements: performance-related descriptions of what a learner at that level can actually do, developed by the Association of Language Testers in Europe (ALTE).

What a learner at each level can do (abridged)

C1Proficient

Understands demanding, longer texts and implicit meaning. Expresses ideas fluently and spontaneously; uses language flexibly for social, academic and professional purposes.

B2Independent

Understands complex text including technical discussion in their field. Interacts with fluency that puts no strain on either party. Produces clear, detailed text and argues a position.

B1Independent

Handles most situations at work, study, and travel. Produces simple connected text on familiar matters; describes experiences and gives reasons for opinions and plans.

A2Basic

Communicates in simple, routine tasks on familiar matters. Describes their background and immediate environment in simple terms.

A1Basic

Understands and uses familiar everyday expressions and basic phrases. Interacts simply if the other person speaks slowly and is prepared to help.

The CEFR proficiency levels as a ladder, listed from the highest down to the lowest — C1 proficient, B2 and B1 independent, then A2 and A1 basic — each rung carrying the abridged can-do statement describing what a learner at that level can actually do.

The learning aims of every STEP skill lesson are derived directly from the can-do statements at the relevant level. For ease of use, STEP subdivides the CEFR bands into a 12-level scale — finer gradations that let a learner (and an employer or placement cell) see movement within a band, not just across bands. The CEFR also specifies the contexts in which language is used — occupational, public, educational — and our items and lessons are written against those contexts, which is why so much of our material is set in offices, interviews, service settings, and workplaces rather than in abstract exercises.

Why insist on an external standard? Because self-referential scores are unfalsifiable. A STEP level means the same thing to a recruiter in Chennai, a placement cell in Coimbatore, and a global capability centre calibrating a hiring bar — and it can be checked against a published international framework rather than taken on our word.

02 · ASSESS

Assessment design

Assessment is where STEP started and remains the anchor of the platform. We conduct over 200,000 assessments a year, and every engagement — enterprise, institutional, or individual — begins with one. Four design principles govern how STEP assessments are built.

Principle 1 — Global standards, externally referenced

Every item, level boundary, and score report maps to the CEFR as described above. We do not grade against a proprietary scale that only we can interpret.

Principle 2 — Calibrated items testing specific skills and sub-skills

Each test item is written to assess a specific skill, at a specific proficiency level, in a specific context. A B2 spoken-production item, for instance, takes the CEFR descriptor — “can deliver announcements on most general topics with clarity, fluency and spontaneity” — and situates it: you work for the railways; a train has been cancelled; prepare the passenger announcement, three minutes to prepare, three minutes to record. The skill, the level, and the context are all deliberate choices, not accidents of the question writer’s imagination.

Items are drawn from a large, actively managed bank: our total pool exceeds 14,000 items, of which roughly 3,100 calibrated items are live in assessments at any time, with 700–800 new items added and calibrated each year based on how items perform with real test-takers. Reading sections, for example, move through five task types — from filling gaps in real correspondence to answering questions on persuasive and argumentative articles — testing sub-skills from reading for gist through reading for inference.

Principle 3 — Adaptive testing

STEP assessments are computer-adaptive: the difficulty of each question adjusts to the test-taker’s performance on the previous ones. This is how a proficiency level can be fixed with high confidence in 15–20 questions per section rather than 60 — the test converges on the learner’s level instead of marching every learner through the same paper. Adaptivity requires every item to be objectively auto-gradeable, which is fully achievable for reading and listening and largely achievable for speaking and writing, with the remainder handled by the open-ended tasks below.

Principle 4 — AI-based evaluation of open-ended speech and writing

The hardest part of language assessment at scale is grading what a learner produces freely. STEP handles this in two stages. In the adaptive portion, structured speaking tasks are evaluated at phoneme level for pronunciation and fluency against expected responses. Then, in the open-ended portion, the learner speaks for 1–2 minutes and writes 100–150 words in response to prompts selected at their tentative level — and the writing is evaluated on four dimensions: ability to stay on topic, readability, grammatical accuracy (error rate against total word count), and lexical range. The open-ended score is combined with the adaptive score to produce the final result. Speech evaluation uses specialist third-party engines alongside our own frameworks; the design principle is that no single automated measure decides a score alone.

Test formats

STEP Plus is the flagship assessment: a comprehensive, proctored evaluation of 60 minutes across 23 tasks, with 15 scored responses in each of listening, speaking, reading, and writing. It is the instrument behind STEP certification and formal benchmarking — the score an employer calibrates a hiring bar against and a learner carries into the job market.

Everything else is a configuration, not a separate product. From the same calibrated item bank and the same CEFR output standard, we build customised assessments to fit the job at hand — a rapid screening pass for high-volume hiring, a short diagnostic to baseline a cohort and place learners into level-appropriate instruction, or a client-specific instrument weighted towards the skills and contexts that matter for a particular role. Duration, task mix, and skill weighting change; the standard behind the score does not. This is deliberate: a screening result, a diagnostic profile, and a certification are comparable because they come off one infrastructure — not three tests that happen to share a brand.

200,000+ assessments a year14,000+ item pool · ~3,100 live700–800 new calibrated items each year15–20 questions to fix a level, adaptively
A STEP score report: an overall score out of 12 beside separate listening, speaking, reading and writing scores, the CEFR level with its can-do statements, lists of what the test-taker performs best in and can do better in, and a CEFR-to-STEP comparison table

Score report · overall score, four skill bands, CEFR level and skills breakdown

A STEP Plus Test Certificate: an overall score out of 12 beside separate listening, speaking, reading and writing scores each mapped to a CEFR level, a STEP-score-to-CEFR-level comparison table, a certificate number, a reference number, the date of examination, and a signature from the Business Head of STEP

STEP Plus certificate · the same CEFR-mapped scale, issued against a certificate number

03 · TEACH

How instruction is structured

Assessment tells us where a learner is; instruction is how they move. STEP delivers instruction three ways — self-paced online lessons, classroom training, and live 1-on-1 coach calls — and combines them in blended programmes. Each is a distinct discipline with its own structure, and each is set out below.

The online lesson

Every productive-skills online lesson follows a four-part arc: interaction → instruction → practice → feedback. The target language is first met inside a natural interaction — a conversation, an email, an announcement — never as a rule on a slide. The instructor then draws attention to the phrases, structure, and tone that did the work in that interaction, and extends them. The learner practises. The learner gets specific feedback on the practice, including why wrong answers were wrong — because the explanation of an error is where the learning happens.

The anatomy of an online lesson

01

Interaction

The target language is first met inside a natural interaction — a conversation, an email, an announcement — never as a rule on a slide.

02

Instruction

The instructor draws attention to the phrases, structure and tone that did the work in that interaction, and extends them.

03

Practice

The learner practises the target language.

04

Feedback

Specific feedback on the practice, including why wrong answers were wrong — because the explanation of an error is where the learning happens.

Tuned per skill

Speaking

Opens with a filmed interaction; closes with the learner producing the function themselves.

Writing

Opens with a model text — what good looks like — before deconstructing its structure, tone and phrasing.

Listening

Sets context first, uses audio from natural situations, and questions across sub-skills.

Reading

Pairs authentic passages with sub-skill questions and feedback that builds reading strategies, not just answers.

The four-part arc every STEP productive-skills lesson follows. One, interaction: the target language is first met inside a natural interaction such as a conversation, an email or an announcement, never as a rule on a slide. Two, instruction: the instructor draws attention to the phrases, structure and tone that did the work, and extends them. Three, practice: the learner practises the target language. Four, feedback: specific feedback including why wrong answers were wrong, because the explanation of an error is where the learning happens. The arc is then tuned per skill — speaking opens on a filmed interaction, writing on a model text, listening sets context first, and reading pairs authentic passages with sub-skill questions.

The arc is tuned per skill. Speaking lessons open with a filmed interaction and close with the learner producing the function themselves. Writing lessons open with a model text — what good looks like — before deconstructing its structure, tone, and phrasing. Listening lessons set context first, use audio from natural situations (announcements, tours, radio), and question across sub-skills. Reading lessons pair authentic passages with sub-skill questions and detailed feedback that builds reading strategies, not just answers. Grammar lessons cover the core competencies with particular attention to the patterns Indian learners most often stumble on — subject–verb agreement, prepositions, articles — surfaced through interaction rather than drilled through rules.

The classroom

Classroom training is where STEP delivers at cohort scale, and it is a distinct discipline from coaching — not a coach call with thirty people in the room. Three design decisions define it.

First, no mixed-ability teaching. Every classroom programme opens with a proctored baseline assessment, and the cohort is grouped into level-based batches — beginner, intermediate, advanced — before the first session. Teaching to a mixed-ability room fails everyone in it: the material is simultaneously too hard for the bottom third and too easy for the top. Batching by measured level is the classroom expression of the same level-based principle that governs our self-paced courses.

Second, activity-based and communicative throughout. A STEP classroom session is built on the CLT progression from controlled practice to freer communicative tasks. The trainer introduces the target function in context and runs controlled practice — structured, scaffolded, high-correction. As competence builds, the scaffolding comes away and the session moves to freer production: role plays, simulated group discussions run against real evaluation rubrics, mock interviews (HR and technical, 1-on-1), presentations to the room. Learners spend the majority of session time producing language, not receiving it. Units are integrated — grammar, reading, listening, speaking, and writing reinforce each other within a unit rather than running as parallel silos — and programmes build to a course-end showcase where learners perform what they can now do, assessed against the same rubrics used throughout.

Third, the classroom never works alone. Contact hours are paired with an individually-assigned online layer — each learner’s self-paced pathway is generated from their own diagnostic, so classroom time reinforces the cohort-level curriculum while the online hours attack each individual’s specific gaps. Printed workbooks extend the classroom beyond contact hours: self-explanatory by design, restating key concepts and providing practice activities a learner can complete with minimal teacher intervention.

Who stands in front of the room matters as much as the method. STEP classroom trainers are CELTA-certified language teaching professionals and industry veterans — people who have run corporate recruitment and client communication themselves — because placement and workplace training requires both credibility about the destination and craft in the teaching.

The coach call

Where the classroom moves a cohort, a coach call moves one person. A live 1-on-1 call applies the same lesson arc to a single learner’s specific gaps: the coach works from the learner’s assessment profile, the record of previous calls, and above all the assignment the learner completed since the last call. The call is never generic. A typical call runs 30 minutes and moves through the four stages below; calls on the IELTS Speaking course and in corporate programmes run 45 minutes, where a wider scope of work needs the extra time. A coach follows the learner, so a call that needs longer on practice takes it from somewhere else:

  1. Feedback on the assignment. Every call ends with an assignment; every next call opens with feedback on it. The coach names what the learner got right and where the errors were — specifically, against the language the learner actually produced. This loop — an assignment set, then feedback on that assignment — is the single most important mechanism in the coaching model.
  2. Mini-teach, where the assignment calls for it. Where the assignment exposes a particular area of weakness — a tense, a register, a pronunciation feature — the coach teaches to it there and then. The stage is conditional: it runs when the learner’s own work has shown something to fix, not on a schedule.
  3. The input session. The coach introduces the topic for that call and the communicative function it turns on — disagreeing with a senior, opening a client call, structuring bad news — inside a realistic scenario drawn from the learner’s own work or study context. Practice sits inside this stage and moves in one direction. Guided practice comes first: the learner produces the language with the coach scaffolding — prompts, corrections at the point of error, second attempts. Freer practice follows: the scaffolding comes away and the learner runs the scenario end-to-end with minimal or no support — a mock announcement, a role-played conversation, a summarised argument — while the coach observes against the four graded signals. What changes between the two is how much support is in the room.
  4. The next assignment. The coach sets a specific assignment shaped by what the call revealed, and assigns the app modules that practise exactly that. The learner does the work before the next call — and the next call opens with feedback on it.

The structure of a coach call

01

Feedback on the assignment

Every call ends with an assignment; every next call opens with feedback on it — what the learner got right, and where the errors were.

02

Mini-teach, if the assignment calls for it

Where the assignment exposes a particular weakness, the coach teaches to it there and then. Conditional, not scheduled.

03

The input session

The call’s topic and target function, inside a realistic scenario. Practice runs here: guided first, with the coach scaffolding, then freer, with minimal or no support.

04

The next assignment

A specific assignment shaped by what the call revealed, plus the app modules that practise exactly that.

A typical call30 min

The four stages of a STEP coach call, in sequence. One, feedback on the assignment the learner completed since the last call — what they got right, and where the errors were. Two, a mini-teach where that assignment exposes a particular area of weakness; this stage is conditional and runs only when the learner’s own work has shown something to fix. Three, the input session, which introduces the call’s topic and communicative function inside a realistic scenario and contains the practice: guided practice first, with the coach scaffolding through prompts and corrections at the point of error, then freer practice, where the scaffolding comes away and the learner works with minimal or no support. Four, the next assignment, shaped by what the call revealed, with the app modules that practise exactly that. A typical call runs 30 minutes; IELTS Speaking and corporate calls run 45.
A learner and coach greeting each other at the start of a live video call
04 · REINFORCE

The blended loop

In blended programmes the modalities are not parallel tracks; they form a loop. The classroom and the coach build skill. AI-powered assessment surfaces each individual’s gaps. Personalised online practice reinforces exactly those gaps between sessions. Endline reassessment confirms — or denies — the uplift. Each element covers what the others cannot: a classroom cannot give thirty learners individual pronunciation feedback; an app cannot notice that a learner’s hesitation comes from confidence rather than vocabulary. The loop can.

The blended loop

01

The classroom and the coach build skill

Contact hours and live 1-on-1 calls move the cohort and the individual.

02

Assessment surfaces the gaps

AI-powered assessment surfaces each individual’s gaps.

03

Online practice reinforces them

Personalised online practice reinforces exactly those gaps between sessions.

04

Reassessment confirms the uplift

Endline reassessment confirms — or denies — the uplift, on the same CEFR-mapped scale.

And the loop runs again

Each element covers what the others cannot. A classroom cannot give thirty learners individual pronunciation feedback; an app cannot notice that a learner’s hesitation comes from confidence rather than vocabulary. The loop can.

The four stages of the STEP blended loop, drawn as a repeating cycle. One, the classroom and the coach build skill through contact hours and live one-to-one calls. Two, AI-powered assessment surfaces each individual’s specific gaps. Three, personalised online practice between sessions reinforces exactly those gaps. Four, endline reassessment on the same CEFR-mapped scale confirms or denies the uplift — and the loop runs again. Each element covers what the others cannot: a classroom cannot give thirty learners individual pronunciation feedback, and an app cannot notice that a learner’s hesitation comes from confidence rather than vocabulary.
05 · MEASURE

What we measure, and what we report

Four signals are graded every cohort, with audio evidence retained for review:

  • Fluency — words per minute, hesitation, recovery after a stumble.
  • Grammar — tense control, register, agreement, and the patterns Indian speakers most often stumble on.
  • Vocabulary — range and precision in workplace, exam, or test-day contexts.
  • Pronunciation — sound by sound, at phoneme level, with retained audio.

Measurement is pre/post by design: a proctored baseline before instruction begins, an endline after it ends, both on the same CEFR-mapped scale. For organisations, this becomes the outcome report — CEFR movement per learner and per cohort, engagement and completion, sub-skill breakdowns, and, where the data is shareable, placement readiness or job-progression deltas. Institutions additionally get a live dashboard with cohort- and individual-level drill-down. The reporting exists for one reason: the person who bought the programme should be able to defend the spend with data, and the learner should be able to defend the credential.

An anonymised STEP client dashboard: total learners, course completion, certification and no-login counts, a red-amber-green learner distribution, and assessment completion rates

Cohort and learner reporting · client dashboard, identities removed

EVIDENCE

The evidence that it works

Methodology claims are cheap; measured outcomes are not. Three kinds of external validation matter to us.

Measured cohort outcomes. In a recent academic-year engagement with a leading South Indian engineering institution — a multi-year partnership now in its fourth year — 1,567 first-year students across 14 branches completed our 50-hour blended programme; 1,527 of them moved up at least one STEP proficiency level between baseline and endline, with 81% completing the online component. Nearly every learner who finished showed measurable uplift on the same externally-referenced scale they started on.

Government and academic selection. STEP’s learning programme was selected under the AICTE/MHRD National Educational Alliance for Technology (NEAT) programme — a selection made specifically on the basis of measurable learner impact and the use of AI tools to drive improvement. Separately, the Tamil Nadu State Council for Higher Education engaged STEP’s Content and Pedagogy team in the redesign of the state’s undergraduate English curriculum — CEFR-aligned, skills-based, and performance-assessed on precisely the principles described on this page.

Institutional and enterprise adoption at scale. Organisations that have used STEP for assessment and training include TCS, Cognizant, L&T, Bank of America, Ford, and ICICI Prudential, and institutions including ISB, Christ University, Manipal Academy, and the Shiv Nadar Foundation — buyers with the sophistication to test claims before renewing.

An anonymised STEP dashboard view comparing diagnostic, practice and certification averages for listening, speaking, reading and writing, with the movement from diagnostic to certification shown for each skill

Measured cohort movement · diagnostic to certification by skill, client dashboard with identities removed

ONGC logo ICICI Prudential logo Wipro logo Cognizant logo Asian Paints logo Tech Mahindra logo Visvesvaraya Technological University logo RV College of Engineering logo Christ University logo Indian Maritime University logo Mohan Babu University logo J.C. Bose University of Science and Technology logo
Organisations that have tested the method — from our corporate and higher-education programmes
07

From the Content and Pedagogy team

English is accessible to everyone. Everyone can learn to speak with confidence.

We have taught in engineering colleges and on contact-centre floors. We have coached vice-presidents and first-year support reps. We have watched students who failed a Class X English board go on, three years later, to train the new joiners.

Everything above — the CLT foundation, the CEFR backbone, the adaptive assessment, the lesson arc, the coach loop — exists in service of three plain observations. People do not need to be told they are smart; they need to be told, specifically, what went wrong in the work they just did. They need someone who will look at the next piece of work and tell them whether it went better. And they need to be taken seriously. The platform sets the practice. The coach reads the work and says what changed. And we treat every learner — engineer, exam aspirant, support rep, study-abroad hopeful — as if the work they are doing is real, because it is.

— STEP CONTENT AND PEDAGOGY TEAM, CHENNAI

Offer Ends in
Loading...