Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

yes. this is just raw implementation of the model arch as described in papers. for complete model training with back propogation we need training pipeline with optmizer and loss calculation.
 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: