Principal Data Scientist Resume
Jupiter, FL
SUMMARY
- Over 20 years of experience in data space such as Data Science, Advanced Data Analytics, and Business Intelligence, with focus on Machine Learning, and Artificial Intelligence using Python (Scikit - Learn, TensorFlow, Spark).
- Education focused in Data Science, Applied Mathematics, and Computer Science.
- Passion for Data Science and Artificial Intelligence. Drive to stay on top of new technologies as they emerge, and never allow my core software engineering skills to atrophy.
- Over 10 years of experience leading Agile software development teams, supervising as many as 65 people, with project budgets up to 7 million dollars. Led multiple diverse development teams to include full-time employees, vendors, and offshore resources.
TECHNICAL SKILLS
Expert Skill Level: Python (TensorFlow, Keras, Scikit-Learn, NumPy, Pandas, SciPy, Matplotlib, OpenCV, Unittest, Gunicorn, Django, and Flask), Hadoop Ecosystems (Cloudera Data Science Workbench, Impala, Hive and Spark) Microsoft SQL, Oracle, Power BI, Tableau, Docker .
Algorithms / Models: Artificial Neural Networks with focus on Recurrent (RNN) and Convolutional Neural Networks (CNN). LSTM, PCA, Auto Encoders, XGBoost, LightGBM, CatBoost, Scikit-learn’s (Linear Regression, Logistic Regression, Support Vector Machines including Kernelized SVMs, Decision Trees, ensembles of trees (Random Forests, Gradient Boosted Trees), K Means Clustering, K-Nearest Neighbors, Naïve Bayes).
AWS: SageMaker, CloudFormation, CLI (Command Line Interface), EC2 (Elastic Cloud Compute), S3 (Simplified Storage Services), ECS (Elastic Container Service), ELB (Elastic Load Balancing), Neptune DB, DynamoDB, Redshift, API Gateway, and other common AWS services.
Azure: HDInsight, Databricks, Data Factory, Azure Machine Learning, Data Science virtual machines.
Advanced Skill Level: Java, C#, C++, JavaScript, HTML, CSS, Bootstrap.
Familiar with: R, Scala.
Version Control: Git, GitHub, TFS, Bitbucket, SVN, Dimensions.
Management Systems / Methodologies: Scrum and Kanban implementations of Agile, CRISP-DM, Waterfall, Hybrid of Agile / Waterfall, Six Sigma, PMBOK.
Agile / Project Management Software: Version One, Jira, MS Project, Smartsheet.
PROFESSIONAL EXPERIENCE
Principal Data Scientist
Confidential - Jupiter, FL
Responsibilities:
- Hands on role in all aspects of critically important Machine Learning / Artificial Intelligence projects that require complex custom models.
- Some project examples include CV solutions for Boston Dynamic’s Spot robotic dog, drone aircraft, thermal imaging, and NLP solutions that were implemented into production at multiple Nuclear Power plants for one of the largest energy companies in the United States.
- I established standards with a heavy focus on automations for how we plan, build, test and deploy Machine Learning Projects. I was able to get buy in from the teams to embrace these standards; subsequently, our productivity, and efficiency significantly increased. This was critical for the teams to become very efficient due to the need for extremely quick turnaround times on projects, and a desire to maintain high levels of quality.
- There was a need for me to be able to learn new technologies very rapidly, then mentor the team on how to use these technologies.
- My work with FPL started with them being my main client when I was Director of Data Science for a consulting company. This was an Artificial Intelligence and Data Science company that specialized in delivering end to end solutions with a focus on (CV) computer vision and (NLP) natural language processing.
- I was able to accept a contract role as Principal Data Scientist with FPL immediately after the consulting company I was with lost the bid for next phase of project with FPL.
.
Senior Artificial Intelligence Engineer
Confidential - Charlotte NC
Responsibilities:
- I ran the AI program at Confidential and led the effort to stand up the company’s Data Science / Artificial Intelligence program from the ground up.
- I built multiple AI applications from end to end. The focus of my work was on deep learning, such as using Deep / Recurrent Neural Networks to solve NLP problems.
- I drove projects from conceptual phase, plan, design, build, test, to deploy.
- I deployed the trained models using Flask and Django REST APIs and built Django web applications to consume the models. I use Docker containers, NGINX, and Gunicorn WSGI to serve the web applications.
- Projects included but not limited to: Chatbot, decision engines that used Massive Multilabel Multiclass Deep Learning Models, Anomaly Detection to detect data quality issues, and many automations.
- Designed a Big Data Solution, using Hortonworks distribution of Hadoop.
- Used distributed computing on Azure cloud to include leveraging CUDA for distributed GPU acceleration in training deep learning models in TensorFlow.
- Advised executives up to the CEO on the AI product backlog.
- My strategy for Data Science projects:
- I automate as much of this process as possible to reduce human error, speed up the process to facilitate exploration of more models, and help find the best model given the business need.
- I look for project candidates that would have a clear value that we could sell to stakeholders.
- Frame the big picture, identify the problem we were trying to solve, articulate the value in non-monetary as well as monetary gains with an estimated dollar amount in value add.
- Gain access to and explore the data looking for distribution, outliers, missing data, and opportunities for feature engineering. Feature engineering can have a greater impact on a model’s usefulness than even the selection of model.
- Process the data.
- When needed: get data to the right shape / format, handle nulls, outliers, feature reduction, feature scaling.
- Short-list most promising models / ensemble of models.
- Create confusion matrix to help with model selection.
- Visualizations on model performance help us communicate our reasons for model selection, especially when models are subject to federal regulation and must be approved by an Enterprise Model Validation Team.
- Fine-tune the system.
- Cross-validation for model validation, test against test data, adjust hyperparameters, adjust features, leverage automated hyperparameter tuning such as grid search / sweep, change models if needed.
- Present the solution.
- This step is critical, because the best solution is of no use if we cannot get approval to implement in production.
- Deploy the solution to production.
- Code hardening steps are verified, test automation, and monitoring are put in place.
Data Scientist / Senior Analyst
Confidential - Charlotte NC
Responsibilities:
- I joined the company as they were standing up a Data Science team, and I took a lead role in shaping the team’s Data Science capabilities.
- The team had difficulty selling the value of the projects and experienced resistance from lines of business who wanted to own their Data Science initiatives.
- I developed a plan and gained approval from stakeholders on a Data Science project within weeks of being at Ally.
- I created a proof of concept and deployed the system to production with it becoming the first productionized Data Science deliverable on the team.
- Created user defined functions (UDFs) in Java to use in Data Science Workbench against the Hadoop data lake.
- Ultimately, we moved away from Java UDFs in favor of doing most of what we do in Python.
- Prototyped solutions in Python, R, and Scala to determine the best language for our team. We unanimously agreed to move forward with Python.
- I implemented numerous process improvements
- I automated tasks for the Data Management Team that they were doing manually.
- I built robust automation framework complete with automated testing, notifications, and error logging.
- I migrated the team to Version One for Agile Project management, as they had no way of managing internal team tasks when I arrived.
- I left this position to take a full-time role.
Technical Lead
Confidential
Responsibilities:
- I discovered ways to solve problems and make improvement using data sources that were not being fully utilized.
- For data analysis, I use the Anaconda distribution for Python 3.6. I use Pandas for pre-processing data such as merging, cleaning, transforming, reshaping and packaging data. SciPy, Scikit-Learn for data modeling leveraging machine learning algorithms. Matplotlib and Microsoft Reporting Services for visualization. I presented findings to Sr. management to drive process improvements.
- Analysis of firewall data exports, integrated with system mapping data, enabled me to make major improvements to the firewall that increased security, reduced chance of human error in future firewall changes, and was structured in a report that went into detailed design documentation of firewalls that had never been captured at that level before in our department.
- Using scripts to analyze server log files, I improved speed and accuracy of root cause analysis of critical systems.
- I utilized exports from our project management system Jira to improve our team’s performance, increased monitoring and control, and improved defect management.
- I identified numerous data governance issues and obtained buy-in from the necessary stakeholders to get them resolved across the enterprise.
- Applications were centered on smart meters that feed data into various downstream systems such as Hadoop, Billing, and some customer facing analytics tools. Custom .NET Applications, 3rd Party applications, Java Applications.
- Led team including the following: Solution Architect, Network Architect, Security Architect, Network Engineers, Oracle / SQL DBAs, Windows Server, Security (Firewall, Access Services), Monitoring (SCOM, CA Wily, ArcSight), .Net Developers, Java Developers, Virtualization Engineers, Test Automation Developers, and Build Automation resources.
- Subject matter expert on Scrum implementation of Agile as teams began transitioning from Waterfall to Agile.
- Performed Scrum Master Role: Led daily standups, backlog refinement, sprint planning, sprint retrospective, and sprint review meetings.
- Other process improvements were the early use of proof of concept prototypes to validate high-risk designs, improved detailed technical designs, implementation plans, as well as vastly expanded the use of automation.
- Led projects implementing security enhancements including VLAN isolation of power controlling systems, jump host solutions including two-factor authentication for direct server access, firewalls (NSX, Checkpoint, and Palo Alto), Citrix GUI solution to control access to sensitive web-based applications, server-based and XML gateway based throttling, and Active Directory domain separation for high-risk assets.
- Led projects implementing Infrastructure enhancements such as leveraged physical to virtual, data center migrations, capacity expansions, migrations to Netscaler load balancing, IPV4 to IPV6 conversions, Disaster Recovery including Active / Active, High Availability Designs. Projects utilized build automation (Puppet, Maven, Urban Code Deploy), test automation, and robust performance/load testing.
Technical Lead \ Project Manager
Confidential - Charlotte NC
Responsibilities:
- Led projects leveraging Big Data solutions that integrated into both internal, and customer-facing systems.
- Data modeling thru conceptual, logical and physical data model’s large complex data sets.
- Enhanced enterprise automation framework using C# and a MS SQL Server back end, that tied into the Microsoft BI stack.
