1. Overview
Implement the complete backend Question & Assessment domain for TOALM V2.
This domain is responsible for creating, validating, storing, retrieving, and evaluating the questions used to assess student understanding.
The system must support the project's intended dynamic question architecture, where controlled question templates and parameters are used to produce different valid questions while preserving the intended topic, learning objective, and difficulty.
The backend must remain the source of truth for question validity and assessment correctness.
Gemma must not be treated as the authority for correct answers, question validity, scoring, or assessment state.
The intended question pipeline is:
Question Template
↓
Parameters
↓
Question Generator
↓
Answer Calculator
↓
Question Validator
↓
Valid Question
↓
Student
This follows the architecture already defined for TOALM V2.
2. Objective
Build a complete backend question and assessment subsystem that can:
- Represent question templates.
- Associate questions with curriculum content.
- Support different question types.
- Represent difficulty.
- Generate controlled question instances.
- Calculate/derive expected answers where applicable.
- Validate generated questions before they reach students.
- Store generated questions.
- Retrieve appropriate questions for a learning session.
- Evaluate student responses.
- Produce a reliable assessment result for the Learning/Attempt domain.
- Provide enough structured information for mastery and adaptive learning later.
- Domain Ownership
The developer responsible for this issue owns the complete backend Question & Assessment domain:
- Question models
- Question-template models
- Generated-question models
- Schemas
- CRUD/repositories
- Question generation logic
- Parameter generation
- Answer calculation
- Validation
- Assessment/evaluation logic
- Question-related services
- API routes
- Database migrations
- Development seed data
- Error handling
- Tests
- API documentation
This issue must result in a usable backend question subsystem, not merely a collection of models.
- Question Architecture
The system must distinguish between:
Question Template
↓
Generated Question
↓
Student Attempt
A template represents the reusable structure.
Example:
Template:
Solve {a}x + {b} = {c}
A generated question may become:
Solve 3x + 6 = 15.
The generated question must contain the information required to evaluate the student's response later.
- Question Template
Implement the question-template structure already represented by:
models/question_template.py
A template should support appropriate information such as:
- Template ID
- Question type
- Subject
- Topic
- Relevant lesson/objective where applicable
- Difficulty
- Template content
- Parameters
- Answer-generation/calculation rules
- Explanation/solution structure
- Validation rules
- Active status
- Created/updated timestamps
The exact fields should follow the project's actual question requirements rather than adding unnecessary generic metadata.
- Curriculum Association
Every usable question must be associated with the curriculum context it assesses.
At minimum, the system should be able to determine:
Subject
↓
Topic
↓
Learning objective / lesson
↓
Question
This is necessary because the Artificial Teacher needs to know what concept the student's response represents.
A question must not become an unclassified piece of content.
- Question Types
The backend should support the question types required by the project's actual educational use cases.
The architecture should allow question types to expand later without rewriting the entire assessment system.
Examples may include:
- Multiple choice
- Short answer
- Numerical answer
- Structured response
Do not build a large examination engine with unnecessary question formats that are outside the current project scope.
The first implementation should prioritize the types required by the Tanzania syllabus learning flow.
- Difficulty
Questions must carry a defined difficulty level.
For example:
Easy
Medium
Hard
Difficulty must be represented as structured backend data rather than being inferred only from the wording of the question.
The difficulty value will later be used by the adaptive learning engine to determine whether the student should receive:
Simpler question
↓
Moderate question
↓
More challenging question
This corresponds to the documented adaptive-learning behavior.
- Parameter Generation
For dynamic questions, implement controlled parameter generation.
Example:
Template:
a × x + b = c
Parameters:
a = 3
b = 6
c = 15
The generated question becomes:
3x + 6 = 15
Parameters must be generated according to the template's constraints so that invalid questions are not produced.
- Answer Calculation
Where the question type permits deterministic calculation, the backend should calculate the expected answer.
Example:
3x + 6 = 15
↓
3x = 9
↓
x = 3
The expected answer should be generated independently of the student's submitted answer.
This is important because the backend must not decide that an answer is correct merely because an AI model says it is correct.
- Question Validation
Every generated question must pass validation before being made available to a student.
Validation should check appropriate conditions such as:
- Required fields exist.
- Curriculum association is valid.
- Difficulty is valid.
- Parameters satisfy constraints.
- Question content is valid.
- Expected answer exists where required.
- Expected answer is consistent with the generated question.
- Question is not structurally malformed.
The intended architecture explicitly requires:
Generator
↓
Validator
↓
Valid Question
rather than blindly allowing generated questions into the system.
- Generated Question
Implement the generated-question functionality represented by:
models/generated_question.py
A generated question should retain enough information to reproduce what the student actually received.
At minimum consider:
- Generated question ID
- Template ID
- Student/session association where appropriate
- Subject/topic
- Question content
- Difficulty
- Parameters if required for reproducibility
- Expected answer
- Explanation/solution where appropriate
- Status
- Creation timestamp
The system must not regenerate a different question when later trying to evaluate an old attempt.
- Question Selection
Implement backend logic for selecting an appropriate question.
The selection logic should be capable of considering:
Student context
↓
Current subject/topic
↓
Learning objective
↓
Required difficulty
↓
Available question templates
↓
Generate/select question
The detailed adaptive decision about why the student needs a particular difficulty belongs to the Adaptive Learning domain.
This issue provides the question mechanism that receives that decision.
- Assessment / Response Evaluation
Implement backend evaluation of supported student answers.
The assessment service should determine:
Student Answer
↓
Compare with expected answer
↓
Assessment Result
Possible result:
Correct
Incorrect
Partially correct
Only implement partial-credit behavior where the chosen question type genuinely supports it.
The evaluation result should be structured so the Attempt domain can store it.
- Assessment Result
Return enough information for the Learning/Attempt domain to process the result.
For example:
{
"question_id": 12,
"result": "correct",
"score": 1.0,
"topic_id": 3
}
The exact API contract should follow the team's agreed conventions.
The Question & Assessment domain should not calculate long-term student mastery.
It only determines the result of the current assessment.
- API
A reasonable API structure is:
GET /questions/{question_id}
POST /questions/generate
POST /questions/{question_id}/validate
POST /assessments/evaluate
Additional endpoints may be introduced where required by the existing project contract.
The endpoints should not expose internal generation details unnecessarily.
- Service Structure
A reasonable implementation could contain:
services/
├── question_service.py
└── assessment_service.py
questions/
├── generator.py
├── parameter_generator.py
├── answer_calculator.py
└── validator.py
The exact filenames can differ, but the responsibilities should remain separated.
The important separation is:
Question Service
↓
Question generation/selection
Assessment Service
↓
Student answer evaluation
Learning Service
↓
Stores attempt and updates learning state later
- Database Migration
Create or update Alembic migrations for the Question & Assessment domain.
The database must support:
Question Template
↓
Generated Question
↓
Curriculum association
Migrations must work from a clean database.
- Seed Data
Provide a small set of development question templates and sample questions.
Seed data should demonstrate the architecture rather than attempting to populate the entire Tanzanian syllabus.
Example:
Mathematics
↓
Linear Equations
↓
Question Template
↓
Easy / Medium / Hard
Do not fill the repository with arbitrary thousands of questions just to demonstrate volume.
- Integration With Other Backend Domains
Curriculum Domain provides:
Subject
Topic
Lesson
Learning Objective
Question Domain provides:
Question Template
Generated Question
Expected Answer
Question Difficulty
Learning Session Domain will provide:
Current Session
Student
Current Learning Context
Attempt Domain will consume:
Question
Student Answer
Assessment Result
Mastery/Adaptive Domain will later consume:
Correctness
Score
Topic
Difficulty
This keeps the responsibilities clean.
- What This Issue Must NOT Implement
Do not implement:
- Mastery calculation
- Weakness detection
- Adaptive decision-making
- Recommendations
- AI teaching responses
- Gemma prompting
- Human teacher functionality
- Frontend functionality
- Full learning-session management
Those belong to subsequent domains.
This issue answers:
«"What question should be presented, is the question valid, and was the student's answer correct?"»
- Call Chain
Learning Request
↓
Question Service
↓
Select Template
↓
Generate Parameters
↓
Generate Question
↓
Calculate Expected Answer
↓
Validate
↓
Store Generated Question
↓
Return Question
For assessment:
Student Answer
↓
Assessment Service
↓
Retrieve Generated Question
↓
Evaluate Answer
↓
Assessment Result
↓
Return Result
The next domain will take that result and update the student's learning state.
- Tests
Test at minimum:
Template
- Template creation
- Valid curriculum association
- Invalid curriculum association
- Difficulty validation
- Template retrieval
Generation
- Valid parameters generated
- Valid question generated
- Invalid parameters rejected
- Generated question retains its expected answer
Validation
- Malformed question rejected
- Invalid answer rejected
- Invalid curriculum association rejected
- Invalid difficulty rejected
Assessment
- Correct answer evaluated correctly
- Incorrect answer evaluated correctly
- Appropriate partial-credit behavior where supported
- Invalid question ID handled correctly
- Assessment cannot evaluate against a different question than the one presented
Integration
- Curriculum → Question Template
- Template → Generated Question
- Generated Question → Assessment
- Definition of Done
This issue is complete when:
- Expected Result
At the end of this issue, the backend should be capable of doing this reliably:
Curriculum
↓
Question Template
↓
Parameters
↓
Generated Question
↓
Validation
↓
Student receives question
↓
Student submits answer
↓
Backend evaluates answer
↓
Assessment Result
That result becomes the input for Issue 05 — Learning Sessions & Student Attempts.
1. Overview
Implement the complete backend Question & Assessment domain for TOALM V2.
This domain is responsible for creating, validating, storing, retrieving, and evaluating the questions used to assess student understanding.
The system must support the project's intended dynamic question architecture, where controlled question templates and parameters are used to produce different valid questions while preserving the intended topic, learning objective, and difficulty.
The backend must remain the source of truth for question validity and assessment correctness.
Gemma must not be treated as the authority for correct answers, question validity, scoring, or assessment state.
The intended question pipeline is:
Question Template
↓
Parameters
↓
Question Generator
↓
Answer Calculator
↓
Question Validator
↓
Valid Question
↓
Student
This follows the architecture already defined for TOALM V2.
2. Objective
Build a complete backend question and assessment subsystem that can:
The developer responsible for this issue owns the complete backend Question & Assessment domain:
This issue must result in a usable backend question subsystem, not merely a collection of models.
The system must distinguish between:
Question Template
↓
Generated Question
↓
Student Attempt
A template represents the reusable structure.
Example:
Template:
Solve {a}x + {b} = {c}
A generated question may become:
Solve 3x + 6 = 15.
The generated question must contain the information required to evaluate the student's response later.
Implement the question-template structure already represented by:
models/question_template.py
A template should support appropriate information such as:
The exact fields should follow the project's actual question requirements rather than adding unnecessary generic metadata.
Every usable question must be associated with the curriculum context it assesses.
At minimum, the system should be able to determine:
Subject
↓
Topic
↓
Learning objective / lesson
↓
Question
This is necessary because the Artificial Teacher needs to know what concept the student's response represents.
A question must not become an unclassified piece of content.
The backend should support the question types required by the project's actual educational use cases.
The architecture should allow question types to expand later without rewriting the entire assessment system.
Examples may include:
Do not build a large examination engine with unnecessary question formats that are outside the current project scope.
The first implementation should prioritize the types required by the Tanzania syllabus learning flow.
Questions must carry a defined difficulty level.
For example:
Easy
Medium
Hard
Difficulty must be represented as structured backend data rather than being inferred only from the wording of the question.
The difficulty value will later be used by the adaptive learning engine to determine whether the student should receive:
Simpler question
↓
Moderate question
↓
More challenging question
This corresponds to the documented adaptive-learning behavior.
For dynamic questions, implement controlled parameter generation.
Example:
Template:
a × x + b = c
Parameters:
a = 3
b = 6
c = 15
The generated question becomes:
3x + 6 = 15
Parameters must be generated according to the template's constraints so that invalid questions are not produced.
Where the question type permits deterministic calculation, the backend should calculate the expected answer.
Example:
3x + 6 = 15
↓
3x = 9
↓
x = 3
The expected answer should be generated independently of the student's submitted answer.
This is important because the backend must not decide that an answer is correct merely because an AI model says it is correct.
Every generated question must pass validation before being made available to a student.
Validation should check appropriate conditions such as:
The intended architecture explicitly requires:
Generator
↓
Validator
↓
Valid Question
rather than blindly allowing generated questions into the system.
Implement the generated-question functionality represented by:
models/generated_question.py
A generated question should retain enough information to reproduce what the student actually received.
At minimum consider:
The system must not regenerate a different question when later trying to evaluate an old attempt.
Implement backend logic for selecting an appropriate question.
The selection logic should be capable of considering:
Student context
↓
Current subject/topic
↓
Learning objective
↓
Required difficulty
↓
Available question templates
↓
Generate/select question
The detailed adaptive decision about why the student needs a particular difficulty belongs to the Adaptive Learning domain.
This issue provides the question mechanism that receives that decision.
Implement backend evaluation of supported student answers.
The assessment service should determine:
Student Answer
↓
Compare with expected answer
↓
Assessment Result
Possible result:
Correct
Incorrect
Partially correct
Only implement partial-credit behavior where the chosen question type genuinely supports it.
The evaluation result should be structured so the Attempt domain can store it.
Return enough information for the Learning/Attempt domain to process the result.
For example:
{
"question_id": 12,
"result": "correct",
"score": 1.0,
"topic_id": 3
}
The exact API contract should follow the team's agreed conventions.
The Question & Assessment domain should not calculate long-term student mastery.
It only determines the result of the current assessment.
A reasonable API structure is:
GET /questions/{question_id}
POST /questions/generate
POST /questions/{question_id}/validate
POST /assessments/evaluate
Additional endpoints may be introduced where required by the existing project contract.
The endpoints should not expose internal generation details unnecessarily.
A reasonable implementation could contain:
services/
├── question_service.py
└── assessment_service.py
questions/
├── generator.py
├── parameter_generator.py
├── answer_calculator.py
└── validator.py
The exact filenames can differ, but the responsibilities should remain separated.
The important separation is:
Question Service
↓
Question generation/selection
Assessment Service
↓
Student answer evaluation
Learning Service
↓
Stores attempt and updates learning state later
Create or update Alembic migrations for the Question & Assessment domain.
The database must support:
Question Template
↓
Generated Question
↓
Curriculum association
Migrations must work from a clean database.
Provide a small set of development question templates and sample questions.
Seed data should demonstrate the architecture rather than attempting to populate the entire Tanzanian syllabus.
Example:
Mathematics
↓
Linear Equations
↓
Question Template
↓
Easy / Medium / Hard
Do not fill the repository with arbitrary thousands of questions just to demonstrate volume.
Curriculum Domain provides:
Subject
Topic
Lesson
Learning Objective
Question Domain provides:
Question Template
Generated Question
Expected Answer
Question Difficulty
Learning Session Domain will provide:
Current Session
Student
Current Learning Context
Attempt Domain will consume:
Question
Student Answer
Assessment Result
Mastery/Adaptive Domain will later consume:
Correctness
Score
Topic
Difficulty
This keeps the responsibilities clean.
Do not implement:
Those belong to subsequent domains.
This issue answers:
«"What question should be presented, is the question valid, and was the student's answer correct?"»
Learning Request
↓
Question Service
↓
Select Template
↓
Generate Parameters
↓
Generate Question
↓
Calculate Expected Answer
↓
Validate
↓
Store Generated Question
↓
Return Question
For assessment:
Student Answer
↓
Assessment Service
↓
Retrieve Generated Question
↓
Evaluate Answer
↓
Assessment Result
↓
Return Result
The next domain will take that result and update the student's learning state.
Test at minimum:
Template
Generation
Validation
Assessment
Integration
This issue is complete when:
At the end of this issue, the backend should be capable of doing this reliably:
Curriculum
↓
Question Template
↓
Parameters
↓
Generated Question
↓
Validation
↓
Student receives question
↓
Student submits answer
↓
Backend evaluates answer
↓
Assessment Result
That result becomes the input for Issue 05 — Learning Sessions & Student Attempts.