-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy path5.py
More file actions
85 lines (61 loc) · 3.49 KB
/
Copy path5.py
File metadata and controls
85 lines (61 loc) · 3.49 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
# -*- coding: utf-8 -*-
"""Job 34 Validate career recommendations.ipynb
Automatically generated by Colab.
Original file is located at
https://colab.research.google.com/drive/1JTz0QznGN0CC6AaP1XFzLhIu6tL5p3vb
Validate career recommendations
Confirm Data Integrity
Before making recommendations:
Check for duplicates: Many entries repeat (e.g. repeated skills for multiple careers).
→ Remove duplicates by Career + Domain + Skill/Sub-skill.
Normalize skill names: e.g. “API Consumption”, “APIs API Design” should be standardized (“API Design”).
Validate thresholds and weights: Each “Threshold” (0–100) pairs with a difficulty level; ensure consistent mappings (e.g. Easy = 40.0, Medium = 60.0, Hard = 75.0).
Define Validation Criteria
You can validate a career recommendation (target career) using one or more of the following feature alignments with; Validation, Metric Description and Example.
Skill overlap (%) - % of skills the user already has vs. expected for the target career. For example; If the Data Engineer role requires 40 skills, user has 28 → 70% readiness
Average difficulty match - Compare user’s level vs. Level Required. For example; If user’s average level < required → recommend learning path before transition
Domain fit - Compare user’s current domains (Programming, Data, Security, etc.) to target. For example; Helps detect career/domain alignment
Tag similarity - Match interest/experience keywords in Tags. For example; E.g. “Streaming, ETL” aligns with Data Engineer but not UX Designer
Threshold-weighted readiness - Weighted scoring using the Weight and Threshold columns for skill importance. For example; Weighted composite = ∑ (skill readiness × weight)
Example Validation Workflow
Let’s say you want to validate if someone should transition to Data Engineer:
Input (User Profile):
Current skills: SQL, Python scripting, CI/CD Basics, AuthN/AuthZ, Unit Testing.
Current Role: Backend Developer
Validation Steps:
1. Filter dataset → Career = 'Data Engineer'
2. Extract target skills and their Level Required.
3. Compute overlap with user’s skills:
* SQL ✓
* Python scripting ✓
* CI/CD ✓
* AuthN/AuthZ ✓
* Unit Testing ✓
→ 5/5 match → readiness = 100%
4. Compute difficulty gap: user skill levels >= required → no gap.
5. Weighted readiness = (∑ weights of matched skills / total career weight) → if > 0.6 → indicate good readiness to transition.
Implement Validation
"""
import pandas as pd
# load dataset
df = pd.read_excel('New jobready_career_skill_dataset_1000_rows_v2.xlsx')
def validate_career(user_skills, target_career):
career_df = df[df['Career'] == target_career]
skills_required = set(career_df['Skill/Sub-skill'].str.lower())
user_normalized = set(s.lower() for s in user_skills)
matched = skills_required & user_normalized
readiness_score = len(matched) / len(skills_required)
threshold_mean = career_df['Threshold (0-100)'].mean()
return {
"career": target_career,
"readiness": round(readiness_score * 100, 2),
"avg_threshold": threshold_mean,
"matched_skills": list(matched)
}
# Example
validate_career(["SQL Basics", "Unit Testing", "CI/CD Basics"], "Data Engineer")
"""Evaluate with Domain Clusters
Use domain grouping (e.g., APIs & Integration, Programming, Quality & Testing) to enhance validation accuracy:
* High overlap in 3+ relevant domains → strong recommendation.
* Low overlap in core domains → suggest bridge learning path before recommending.
"""