Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions macros/remove_deprecated_models.sql
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@
("mv", reporting, "video_transcript_events"),
("mv", reporting, "fact_learner_course_grade"),
("mv", reporting, "fact_learner_course_status"),
("mv", reporting, "dim_subsection_performance"),
("mv", event_sink, "most_recent_course_blocks"),
("mv", reporting, "fact_enrollment_status"),
("mv", event_sink, "most_recent_object_tags"),
Expand Down
2 changes: 0 additions & 2 deletions models/problems/dim_problem_responses.sql
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,6 @@ with
select
org,
course_key,
emission_time,
block_id_short,
response,
success,
Expand All @@ -71,7 +70,6 @@ from final_results
group by
org,
course_key,
emission_time,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm trying to understand how this is causing the replacing merge tree issues, were there events with duplicate timestamps being incorrectly aggregated here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the timestamps were causing events to NOT be aggregated but we need them to. the mv is keyed on response and the response_count should have been updated to the aggregate each time a new event came in, but the emission_time made the count always 1 and then the mv would just replace the previous record with the same values and still a count of 1

block_id_short,
response,
success,
Expand Down
46 changes: 11 additions & 35 deletions models/problems/dim_problem_results.sql
Original file line number Diff line number Diff line change
Expand Up @@ -8,52 +8,28 @@
}}

with
first_success as (
select
org,
course_key,
object_id,
argMin(attempts, emission_time) as attempts,
success,
actor_id
from {{ ref("problem_events") }}
where verb_id = 'https://w3id.org/xapi/acrossx/verbs/evaluated' and success
group by org, course_key, object_id, actor_id, success
),
events as (
select distinct org, course_key, object_id, problem_id
from {{ ref("problem_events") }} events
where verb_id = 'https://w3id.org/xapi/acrossx/verbs/evaluated'
),
final_results as (
select
events.org as org,
events.course_key as course_key,
first_success.success as success,
first_success.attempts as attempts,
first_success.actor_id as actor_id,
splitByChar('@', events.problem_id)[3] as block_id_short,
last_response.org as org,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've been out of this for a while, can you write up a quick explanation of the fix?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the original query was taking the first successful response for each actor and left joining all attempts - which means that if an actor never had a successful attempt, their actor_id and number of attempts would be NULL. we then did a distinct at the end which would only keep 1 record for each problem that never had a correct attempt instead of actually counting how many attempts were made

the new query uses the last response for each actor regardless of if its successful or not. this way, we can get an accurate count of incorrect and correct responses and all data is populated for each attempt.

last_response.course_key as course_key,
last_response.success as success,
last_response.attempts as attempts,
last_response.actor_id as actor_id,
splitByChar('@', last_response.problem_id)[3] as block_id_short,
{{
format_problem_number_location(
"events.object_id", "blocks.display_name_with_location"
"last_response.object_id", "blocks.display_name_with_location"
)
}}
from events
from {{ ref("dim_learner_last_response") }} last_response
left join
{{ ref("dim_course_blocks") }} blocks
on (
events.course_key = blocks.course_key
and events.problem_id = blocks.block_id
)
left join
first_success
on (
first_success.org = events.org
and first_success.course_key = events.course_key
and first_success.object_id = events.object_id
last_response.course_key = blocks.course_key
and last_response.problem_id = blocks.block_id
)
)
select distinct
select
org,
course_key,
success,
Expand Down
51 changes: 0 additions & 51 deletions models/problems/dim_subsection_performance.sql

This file was deleted.

21 changes: 0 additions & 21 deletions models/problems/schema.yml
Original file line number Diff line number Diff line change
Expand Up @@ -237,27 +237,6 @@ models:
data_type: String
description: '{{ doc("email") }}'

- name: dim_subsection_performance
columns:
- name: org
data_type: String
description: '{{ doc("org") }}'
- name: course_key
data_type: String
description: '{{ doc("course_key") }}'
- name: block_id
data_type: String
description: '{{ doc("block_id") }}'
- name: total_avg
data_type: Float32
description: Average scaled score per course/problem
- name: score_range
data_type: String
description: Score range
- name: score_range_count
data_type: Int32
description: Count of average scores in each range bucket

- name: dim_problem_coursewide_avg
columns:
- name: org
Expand Down
4 changes: 2 additions & 2 deletions models/problems/unit_tests.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -40,10 +40,10 @@ unit_tests:
config:
tags: 'ci'
given:
- input: ref('problem_events')
- input: ref('dim_learner_last_response')
format: sql
rows: |
select * from problem_events
select * from dim_learner_last_response
- input: ref('dim_course_blocks')
format: sql
rows: |
Expand Down
Loading