Sr Data Engineer Resume
San Diego, CA
SUMMARY
- PhD of Computational Chemistry with 15+ years’ experience of professional IT and academic experience which includes 4+ years of experience in Hadoop, Spark ecosystem related technologies.
- 15+ years’ experience of Unix/Linux shell scripting, Java, Perl, Python, C/C#/C++, and SQL programing on Linux, Mac OS X and Windows.
- Experience in software development life cycle, business requirement analysis/design, programming, data warehousing and business intelligence concepts, and support of systems application architecture.
- Excellent understanding / knowledge of Hadoop architecture and various components such as Hive, HDFS, Map/Reduce, Pig, Hbase, Spark, Oozie, Flume/Kafka, Presto, etc.
- Extensive hold over Hive/Pig core functionality by writing custom java UDFs.
- Significant experience on developing real - time ETL processes to load data from multiple data sources to HDFS using Flume/Kafka/PigLatin custom UDFs, perform structural modifications using MapReduce, Hive and analyze data using visualization/reporting tools.
- Strong experience in designing data pipelines to process high transaction volumes and large datasets.
- Strong knowledge on Hadoop HDFS architecture and MapReduce framework with the MapR distributions of Apache Hadoop.
- Hands on experience in IDE tools like Eclipse, IntelliJ, Visual Studio.
- Experiencedin software configuration management usingSVN, GIT, Maven and Jenkins.
- Solid mathematical background and experience on data mining, machine learning and numerical computation with Hadoop, MatLab and R.
- Strong team player with great verbal, written communication and interpersonal skills.
- Excellent problem solving and communication skills, ability to work within and across diverse teams.
TECHNICAL SKILLS
Big Data: MapR, Hadoop, MapReduce, HDFS, Hive, Pig, Kafka, Oozie, Presto, Spark
Database: SQL Server, Oracle, MySQL, HP Vertica
Reporting: Excel, SSRS, Tableau, MicroStrategy, OBIEE, QlikView, Sharepoint 2010/2013
Tools: Eclipse, IntelliJ, Visual Studio, Wherescape RED, Toad, ER/Studio, SQL*plus
Data warehousing Methodologies/Processes: Dimensional Modeling (Star/Snowflake Schema Modeling, Fact and Dimension Tables), Kimball Methodology, Inmon Methodology, Maintenance of Operational Data Store (ODS) and Enterprise Data warehouse (EDW), OLAP, Metadata Management, Data Migration and Data Cleansing Techniques, Data Profiling, Data Quality and Data Validation Scripts
PROFESSIONAL EXPERIENCE
Sr Data Engineer
Confidential, San Diego, CA
Responsibilities:
- Assisting customers in demonstrating value from a Hadoop environment to drive business forward. Providing guidance on Big-Data solution designs, use case discovery, process development and machine-learning algorithms customized within the Hadoop ecosystem.
- Responsible for maintaining, upgrading and monitoring Hadoop clusters for internal customers.
- Trained data scientists to use MapReduce, Pig/Hive, Oozie on Hadoop to build of models to predict credit and identity fraud risk with large amounts of consumers’ credit behavior data.
- Developed Pig/Hive UDFs with Java and wrap with Maven. Coding MapReduce programs to parse the raw data, encrypt, clean and store the refined data in partitioned tables in EDW.
- Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with EDW reference tables and historical metrics.
- Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System and Pig to pre-process the data.
- Experiencedin software configuration management usingSVN, Git, Maven and Jenkins.
- Recompiled Presto for MapR File System and developed Presto UDFs with Java and customized presto JDBC drivers to support inserting back into relational database tables.
- Set up Kafka nodes to stream data from MySQL to Hive tables.
- Designed Drools/Rules Engine to store and read rules from MySQL tables; design nested JSON data structures for source records.
- Developed schemas (JSON, AVRO, etc.) and related data structures for Pig UDFs and output data storage.
- Designed data pipelines to process terabytes of encrypted data sets daily using MapReduce JAVA, MapR/Hadoop, Maven, JSON, Pig and HiveQL.
- Loaded data into the cluster from dynamically generated files using Flume and from relational database management systems using Sqoop.
- Worked within a Linux/Unix environment; sFTP, GPG encryption/decryption and bash scripting for full automation of functions.
- Integration with MySQL, Sqoop, YARN, ZooKeeper, Warden and other file management tools.
- Developed self-service BI capabilities to automate the production and delivery of reports and analyses in support of various departments and business groups in achieving their business plans.
Sr Data Warehouse Engineer
Confidential, Tallahassee, FL
Responsibilities:
- Extracted files from RDBMS through Sqoop and placed in HDFS and run Hadoop streaming jobs to process terabytes of data; Supported Map/Reduce Programs those are running on the cluster.
- Involved in loading data from UNIX file system to HDFS and creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
- Imported/exported files to and from EDW and Hadoop, create SSIS packages to move data from Hadoop to relational database, use the Hive ODBC driver to create a Linked Server from SQL Server to Hive and build SSAS cubes, configure the Hive ODBC driver to access the Hadoop cluster to create PowerPivot and PowerView reports in MS Excel and on SharePoint sites.
- Extracted data from Hive using SSIS, and used them as additional dimensions for SSAS tabular model; prepare the detailed level ETL mapping specification documents describing the algorithms and flowchart and mentioning the source systems, change data capture logic, transformation logic involved in each field, the lookup tables, lookup logic used and also the target systems per each individual mapping involved with the Data Mart.
- Executed queries using Hive and developed MapReduce jobs to analyze data, developed Pig Latin scripts to extract the data from the web server output files to load into HDFS, developed the Pig UDF's to preprocess the data for analysis.
- Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with EDW reference tables and historical metrics.
- Experience using modeling tools, logical and physical database design using ER/Studio and Wherescape RED; Maintaining and updating database models, generating or modifying the database schemas Data Mapping and DWH loading.
- Responsible for Database support, troubleshooting, planning and migration. Resource planning and coordination for application migrations with project managers, application and web app teams. Project involved guidance and adherence to standardized procedures for planned data center consolidation for worldwide centers using in-house corporate and third party applications based on SQL 2005 in upgrade project to SQL 2008/R2.
- Designed and developed SSIS Packages to import and export data from XML, MS Excel, and SQL Server 2008/R2/2012.
- Migrated over 1,500 SSRS reports to SharePoint 2013 website using PowerShell script.
- Designed and implemented Dashboards and Scorecards with business KPIs using Performance Point Server and published them Via Microsoft Office Share point Server (MOSS).
Staff Scientist
Confidential, CA
Responsibilities:
- Developed fast parallel computing algorithms in linux cluster environments with hundreds of nodes.
- Developed drug candidate database to facilitate the data management and data mining for natural product drug discovery efforts using TOAD and MYSQL.
- Teached software design courses in NBCR summer school for hundreds of graduate students and professors.
- Developed the C++ software package SMOL with machine learning, statistical signal processing on multiscale biological signaling transduction for computational chemistry and neuroscience, for example: acetylcholine diffusion in a single acetylcholinesterase (AChE) cluster, APS2- channeling in sulfate activating complexes, synaptic transmission and drug-induced modification of ionic conductance in neuromuscular junction, Ca2+ diffusion in cardiovascular t-tubule in the normal and failing rodent heart.
- Programmed the hybrid Brownian dynamics/finite element code to implement multi-scale diffusion studies.
- Coded a Python tool chain for tetrahedral mesh generation, optimization and refinement and visualization of continuum modeling output with TetGen, OpenDX, GMV, etc.
- Coded a PDB to PQR convertor based on Amber xLeap with C++.
Associate Research Scientist
Confidential
Responsibilities:
- Applied atomic multipole machine learning on different organic molecules and radicals to explore various chemical and physiological properties. The project concentrates on visualizing and drawing 3-D target molecules, analyzing the energy profile of bond breaking/forming and pi-cation interactions, and comparing various functioning groups, anions or cations.
REFERENCES Available upon request
