We provide IT Staff Augmentation Services!

Data Modeler Resume

0/5 (Submit Your Rating)

NY

SUMMARY

  • Over 8 years of experience working as a Data Analyst / Data Analytics Engineer.
  • Experienced on Big data Hadoop, Spark, business intelligence (and BI technologies)
  • Experienced in employing Pyspark, SQL, Scala to work on building batch jobs, scalable applications and Fraud defenses.
  • Adept at using SAS Enterprise suite, R, Python, Scala and Big Data related technologies including Hadoop, Hive, Pig, Sqoop, Cassandra, Spark, Oozie, Flume, Map - Reduce and Cloudera Manager for design of business intelligence applications
  • Hands on experience with Machine Learning, Regression Analysis, Clustering, Boosting, Classification, Principal Component Analysis and Data Visualization Tools
  • Familiarity with Crystal Reports, and SSRS - Query, Reporting, Analysis and Enterprise Information Management
  • Excellent knowledge on creating reports on Pentaho Business Intelligence.
  • Experienced in Database using Oracle, XML, DB2, Teradata 15/14, Netezza, server, Postgre
  • Experienced with machine learning tools and libraries such as Python-Scikit-learn, R, Spark and Weka

TECHNICAL SKILLS

Data Modelling Tools: Erwin r9.6, 9.5, 9.1, 8.x, Rational Rose, ER/Studio, MS Visio

Programming Languages: Oracle PL/SQL, Julia, UNIX shell scripting, Java. Python 3.X(NumPy, SciPy, Pandas,pyspark), R (Caret, Weka, ggplot)

Big Data Technologies: Hadoop( Hive, HDFS, MapReduce, Pig, Kafka), SPARK 2.2.4, Cassandra, MongoDB, AWS(RDS, Dynamodb, Redshift)

Reporting Tools: Crystal reports XI, Business Intelligence, SSRS, Business Objects 5.x/ 6.x

AWS Cloud: EC2, S3, RDS, Dynamodb, Kinesis, Redshift

ETL: Informatica Power Centre, SSIS.

Project Execution Methodologies: Ralph Kimball and Bill Inmon, datawarehousing methodology, Rational Unified Process (RUP), Rapid Application Development (RAD), Joint Application Development (JAD)

Operating Systems: Windows, UNIX, LINUX, Mac OS

Databases: Oracle /11g/12c/10g/9i, SQL Server, MySQL, PostgreSQL

PROFESSIONAL EXPERIENCE

Confidential, Malvern, PA

Data Analyst/AWS Developer

Responsibilities:

  • Develop and add features to existing applications built with Spark on a Scala development platform on the data and services with AWS cloud.
  • Worked on Writing unit test cases for the applications to verify the credibility of the data on AWS.
  • The results of these application unit tests were recorded in JSON, which was sent to hundreds of business owners across the organization through AWS SNS.
  • Worked on AWS services like Elastic Map Reduce EMR, S3, Service catalog etc.
  • Used HUE to validate and develop data applications with HIVE and Presto.
  • Worked on query and transformation clusters in EMR stack.
  • Worked extensively on production environment with STS, GIT and Bamboo build and release tool.
  • Worked on AGILE environment using JIRA and user stories.
  • Wrote extensive SQL queries on HIVE and IBM DB2 for creating data applications and validations of data.
  • Worked with Control-M tool to start, monitor and schedule jobs as integrated with AWS EMR in ENG, Test and Production.

Environment: Spark, AWS EMR, AWS S3, STS, Scala, Spark-SQL, SQL IBM DB2, Control-M, Bamboo, GIT, AWS Service Catalog, AWS SNS

Confidential, Richmond VA

Sr. Data Analyst / Data scientist

Responsibilities:

  • Develop Fraud Batch Defenses On using Spark and AWS. Using Python and Spark - SQL to build stable Defenses against fraud.
  • This involves working on identifying the fields in legacy Teradata database and crosswalk mapped fields which are migrated into Cloud, stored in Onelake-S3.
  • Coordinated with different teams across Confidential, while involving multiple work streams
  • The defenses worked saves millions of dollars for Confidential against online Fraud.
  • Experienced working with all kinds of source formats, which includes - CSV, JSON, AVRO, Parquet, etc.
  • Experienced working on data with billions of records, coming from disparate source engines, namely - AWS S3, Onelake-S3, Cerebra, and Snowflake.
  • Identified crosswalk between the Teradata fields and corresponding One lake S3 tables on Nebula for development stages.
  • Used mount functionality in spark to get access to the One Lake datasets for consumption purposes using control plane with particular, ASV, BAPCI and LOB ARN IAM role for a given databricks NPI cluster.
  • Experienced working with different types of source, namely, application data, process data, TSYS data, chordiant case data.
  • Worked on catapulting data from teradata to snowflake to consume on Databricks.
  • Implemented the batch defenses on a framework, which works on AIDE (Automated Investigations Decisioning engine). which was developed using pyspark and python API.
  • Written extensive Pyspark SQL queries for the defense and validated the queries as comparing to the results on the legacy teradata platform and snowflake cloud accordingly.

Environment: Onelake-S3, Aws S3, Snowflake, Databricks, Spark, Pyspark, Spark-SQL, Python, SQL, Teradata BTEQ.

Confidential, Reston VA

Data Analytics Engineer/ Data scientist

Responsibilities:

  • Performed scoring and financial forecasting for collection priorities using Python, R and SAS machine learning algorithms.
  • Builtdatapipelines for reporting, alerting, anddatamining. Experienced with table design anddata management using HDFS, Hive, Impala, Sqoop, MySQL, and Kafka.
  • Developed Python modules for machine learning & predictive analytics on AWS. Implemented a Python-based distributed random forest via Python streaming.
  • Worked with statistical models for data analysis, predictive modelling, machine learning approaches and recommendation and optimization algorithms.
  • Working in Business and Data Analysis, Data Profiling, Data Migration, Data Integration and Metadata Management Services.
  • Worked extensively on Databases preferably Oracle 11g/12c and writing PL/SQL scripts for multiple purposes.
  • Built models using Statistical techniques like Bayesian HMM and Machine Learning classification models like XG Boost, SVM, and Random Forest using R and Python packages.
  • Worked with data compliance teams, data governance team to maintain data models, Metadata, data Dictionaries, define source fields and its definitions.
  • Worked with BigDataTechnologies such Hadoop, Hive, MapReduce
  • Setup storage and data analysis tools in Amazon Web Services cloud computing infrastructure.
  • A highly immersive Data Science program involving Data Manipulation & Visualization, Web Scraping, Machine Learning, Python programming, SQL, GIT, Unix Commands, NoSQL, MongoDB, Hadoop.
  • Handled importingdatafrom variousdatasources, performed transformations using Hive, MapReduce, and loadeddatainto HDFS
  • Managed existing team members, lead the recruiting and onboarding of a larger Data Science team that addresses analytical knowledge requirements.
  • Worked directly with upper executives to define requirements of scoring models.
  • Developed a model for predicting a debtor setting up a repayment rehabilitation program for student loan debt.
  • Developed a model for predicting repayment of debt owed to small and medium enterprise (SME) businesses.
  • Developed a generic model for predicting repayment of debt owed in the healthcare, large commercial, and government sectors.
  • Created SQL scripts and analyzed the data in MS Access/Excel and Worked on SQL and SAS script mapping.

Environment: R, SQL, Python 2.7.x, SQL Server 2014, regression, logistic regression, random forest, neural networks, Topic Modeling, NLTK, SVM (Support Vector Machine), JSON, XML, HIVE, HADOOP, PIG, Sklearn, SciPy, GraphLab, No SQL, SAS, SPSS, Spark, Hadoop, Kafka, H Base, MLib.

Confidential, Alexandria, VA

Data scientist/ R Developer

Responsibilities:

  • Designed an Industry standard data Model specific to the company with group insurance offerings
  • Translated the business requirements into detailed production level using Workflow Diagrams, Sequence Diagrams, Activity Diagrams and Use Case Modeling
  • Conceptualized the most-used product module (Research Center) after building a business case for approval, gathering requirements and designing the User Interface
  • A team member of Analytical Group and assisted in designing and development of statistical models for the end clients.
  • Coordinated with end users for designing and implementation of e-commerce analytics solutions as per project proposals.
  • Conducted market research for client; developed and designed sampling methodologies, and analyzed the survey data for pricing and availability of clients' products.
  • Investigated product feasibility by performing analyses that include market sizing, competitive analysis and positioning.
  • Successfully optimized codes in Python to solve a variety of purposes in data mining and machine learning in Python.
  • Building programing logics for developing analysis datasets by integrating with variousdatamarts in the sandbox environment
  • Facilitated stakeholder meetings and sprint reviews to drive project completion.
  • Successfully managed projects using Agile development methodology
  • Project experience in Data mining, segmentation analysis, business forecasting and association rule mining using Large Data Sets with Machine Learning.
  • Automated Diagnosis of Blood Loss during Accidents and Applied Machine Learning algorithms to diagnose blood loss from vital signs (ECG, HF, GSR, etc.).
  • Demonstrated performances of 94.6% on par with state-of-the-art models used in industry

Environment: R, Windows XP/NT/2000, SQL Server 2005/2008, SQL, Oracle8i/10g, DB2, MS Excel, Mainframes MS Visio, Crystal Reports 9., Python, R Studio, Shiny, Excel 2013.

Confidential, NY

Data Modeler

Responsibilities:

  • Responsible for technical data governance, enterprise wide data modeling and database design.
  • Used Model Mart of ERwin for effective model management of sharing, dividing and reusing model information and design for productivity improvement.
  • Conducted detailed and comprehensive Business Analysis by working with the IT staff, Business Staff, SME's, and other stakeholders to identify the system, operational requirements and process improvements.
  • Designed an Industry standard data Model specific to the company with group insurance offerings
  • Translated the business requirements into detailed production level using Workflow Diagrams, Sequence Diagrams, Activity Diagrams and Use Case Modeling
  • Worked with Development DBA to assist and support developers with SQL performance tuning, query tuning and code reviews
  • Analyzed business requirements, system requirements, data mapping requirement specifications, and responsible for documenting functional requirements and supplementary requirements in Quality Center.
  • Extensively worked on Source to Target mapping for business need and documentation purposes.
Environment: Erwin r8, Informatica, Windows XP/NT/2000, SQL Server 2005/2008, SQL, Oracle8i/10g, DB2, MS Excel, Mainframes MS Visio, Rational Rose, Requisite Pro.

Confidential

Data Modeler

Responsibilities:

  • Conducted one-to-one sessions with business users to gather data for Data Warehouse requirements.
  • Part of team analyzing database requirements in detail with the project stakeholders through Joint Requirements Development(JRD) sessions.
  • Developed an Object modeling in UML for Conceptual Data Model using Enterprise Architect.
  • Developed logical and Physical data models using ERwin to design OLTP system for different applications.
  • Facilitated transition of logical data models into the physical database design and recommended technical approaches for good data management practices.
  • Performed numerous SQL Statements to load and extractdatato and from database for testing.
  • Worked with SQL Server Integration Services in extracting data from several source systems and transforming the data and loading it into ODS.
  • Worked with DBA group to create Best-Fit Physical Data Model from the Logical Data Model using Forward engineering using ERWin
Environment: Erwin r9.6, DB2, Teradata, SQL-Server 2008, Informatica 8.1, Enterprise Architect, Power Designer, MS SSAS, Crystal Reports, SSRS, ER Studio, Lotus Notes, Windows XP, MS Excel, word and Access.

We'd love your feedback!