-
Notifications
You must be signed in to change notification settings - Fork 10
Aspects data fixes #170
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Aspects data fixes #170
Changes from all commits
7a62923
30db978
ae6b8e0
fe35f27
5a16081
9ed72bf
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -8,52 +8,28 @@ | |
| }} | ||
|
|
||
| with | ||
| first_success as ( | ||
| select | ||
| org, | ||
| course_key, | ||
| object_id, | ||
| argMin(attempts, emission_time) as attempts, | ||
| success, | ||
| actor_id | ||
| from {{ ref("problem_events") }} | ||
| where verb_id = 'https://w3id.org/xapi/acrossx/verbs/evaluated' and success | ||
| group by org, course_key, object_id, actor_id, success | ||
| ), | ||
| events as ( | ||
| select distinct org, course_key, object_id, problem_id | ||
| from {{ ref("problem_events") }} events | ||
| where verb_id = 'https://w3id.org/xapi/acrossx/verbs/evaluated' | ||
| ), | ||
| final_results as ( | ||
| select | ||
| events.org as org, | ||
| events.course_key as course_key, | ||
| first_success.success as success, | ||
| first_success.attempts as attempts, | ||
| first_success.actor_id as actor_id, | ||
| splitByChar('@', events.problem_id)[3] as block_id_short, | ||
| last_response.org as org, | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I've been out of this for a while, can you write up a quick explanation of the fix?
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. the original query was taking the first successful response for each actor and left joining all attempts - which means that if an actor never had a successful attempt, their actor_id and number of attempts would be NULL. we then did a distinct at the end which would only keep 1 record for each problem that never had a correct attempt instead of actually counting how many attempts were made the new query uses the last response for each actor regardless of if its successful or not. this way, we can get an accurate count of incorrect and correct responses and all data is populated for each attempt. |
||
| last_response.course_key as course_key, | ||
| last_response.success as success, | ||
| last_response.attempts as attempts, | ||
| last_response.actor_id as actor_id, | ||
| splitByChar('@', last_response.problem_id)[3] as block_id_short, | ||
| {{ | ||
| format_problem_number_location( | ||
| "events.object_id", "blocks.display_name_with_location" | ||
| "last_response.object_id", "blocks.display_name_with_location" | ||
| ) | ||
| }} | ||
| from events | ||
| from {{ ref("dim_learner_last_response") }} last_response | ||
| left join | ||
| {{ ref("dim_course_blocks") }} blocks | ||
| on ( | ||
| events.course_key = blocks.course_key | ||
| and events.problem_id = blocks.block_id | ||
| ) | ||
| left join | ||
| first_success | ||
| on ( | ||
| first_success.org = events.org | ||
| and first_success.course_key = events.course_key | ||
| and first_success.object_id = events.object_id | ||
| last_response.course_key = blocks.course_key | ||
| and last_response.problem_id = blocks.block_id | ||
| ) | ||
| ) | ||
| select distinct | ||
| select | ||
| org, | ||
| course_key, | ||
| success, | ||
|
|
||
This file was deleted.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
I'm trying to understand how this is causing the replacing merge tree issues, were there events with duplicate timestamps being incorrectly aggregated here?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
the timestamps were causing events to NOT be aggregated but we need them to. the mv is keyed on response and the response_count should have been updated to the aggregate each time a new event came in, but the emission_time made the count always 1 and then the mv would just replace the previous record with the same values and still a count of 1