The switch from pickle to json for model serialization/deserialization was causing ModelVersion objects to be incorrectly reconstructed:
-
Dataclass Reconstruction Issue: When
ModelVersiondataclass instances were serialized to JSON usingdefault=str, they became plain dictionaries upon deserialization instead of properModelVersionobjects. -
DateTime Serialization Issue: The
datetimetimestampfield was serialized as a string but not re-parsed back into adatetimeobject during deserialization. -
AttributeError Results: This led to
AttributeErrorwhen trying to:- Access
ModelVersionattributes like.version_id,.timestamp - Call
.isoformat()on the timestamp field
- Access
llm/continuous_learning_system.py:281-309(rollback_model())llm/continuous_learning_system.py:559-584(_create_model_version())llm/continuous_learning_system.py:626-670(_load_or_create_model())
Added two new classes to handle proper serialization/deserialization:
class ModelVersionJSONEncoder(json.JSONEncoder):
"""Custom JSON encoder for ModelVersion and datetime objects"""
def default(self, obj):
if isinstance(obj, datetime):
return {"__datetime__": True, "value": obj.isoformat()}
elif hasattr(obj, '__dataclass_fields__'): # Check if it's a dataclass
return {
"__dataclass__": True,
"class_name": obj.__class__.__name__,
"data": asdict(obj)
}
return super().default(obj)
class ModelVersionJSONDecoder(json.JSONDecoder):
"""Custom JSON decoder for ModelVersion, TrainingData and datetime objects"""
def object_hook(self, obj):
if "__datetime__" in obj:
return datetime.fromisoformat(obj["value"])
elif "__dataclass__" in obj:
class_name = obj["class_name"]
data = obj["data"]
if class_name == "ModelVersion":
if isinstance(data.get("timestamp"), str):
data["timestamp"] = datetime.fromisoformat(data["timestamp"])
return ModelVersion(**data)
elif class_name == "TrainingData":
if isinstance(data.get("timestamp"), str):
data["timestamp"] = datetime.fromisoformat(data["timestamp"])
return TrainingData(**data)
return objBefore (buggy approach):
# Using pickle
with open(version_path, "rb") as f:
model_data = pickle.load(f)
# Or using naive JSON with default=str (causes the bug)
json.dumps(model_data, default=str) # Converts objects to stringsAfter (fixed approach):
# Save with custom JSON encoder
with open(version.file_path, "w", encoding="utf-8") as f:
json.dump(model_data, f, cls=ModelVersionJSONEncoder, indent=2)
# Load with custom JSON decoder
with open(json_path, "r", encoding="utf-8") as f:
model_data = json.load(f, cls=ModelVersionJSONDecoder)The fix maintains backward compatibility by:
- Checking for both
.jsonand.pklfiles - Falling back to pickle loading if JSON files aren't found
- Prioritizing JSON files over pickle files for new saves
rollback_model(): Now tries JSON first, falls back to pickle_create_model_version(): Saves models as JSON using custom encoder_load_or_create_model(): Loads JSON files preferentially, falls back to pickle
The fix was verified to ensure:
- ✅ ModelVersion objects are properly reconstructed from JSON
- ✅ datetime fields are correctly parsed back to datetime objects
- ✅
timestamp.isoformat()method calls work correctly - ✅ No AttributeError when accessing ModelVersion attributes
- ✅ Backward compatibility with existing pickle files
- Human-Readable Storage: JSON files are easier to inspect and debug
- Cross-Platform Compatibility: JSON is more portable than pickle
- Type Safety: Custom decoder ensures proper object reconstruction
- Datetime Preservation: Proper handling of datetime serialization/deserialization
The bug has been fully resolved while maintaining backward compatibility with existing pickle-based model files.