Sr Data Engineer Resume
VirginiA
SUMMARY
- 11+ years of overall experience with Big Data systems and Java/J2EE applications.
- 6+ years of Hadoop/Big Data Technologies like Spark, MapReduce, Hive, HBase, sqoop, Solr, oozie, Pig and zookeeper.
- 4 + years of experience in Cloud technologies (AWS) like EC2, EMR, S3, Lambda, cloud watch, etc
- Experience in DevelopingSparkapplications usingSpark Core, Spark SQL and Data Frames/Data Sets/RDD API for data extraction, transformation and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns.
- Expertise in developing MapReduce, Spark jobs to facilitate the flexibility of ETL.
- Hands on experience with performance optimization techniques for data processing in Hive, Spark, Pig & Map - Reduce.
- Experience in implementing Data Quality metrics and Data reconcile using spark and SQL.
- Experience in creating solr/Elasticsearch collections, loading data using spark, writing solr/Elasticsearch queries to enable the fast searching on data.
- Experience with writing complex SQL queries in IBM DB2, Oracle, Sybase, MS SQL Server 2005/2008 and Netezza
- Using Sqoop for import multiple table data from RDBMS to Hadoop environment.
- Using Pig to aggressively analyze and expose the various facts of the data on the fly.
- Ability to fine tune application code to maximize its potential in minimal time for SPARK and HIVE.
- Experience Working with Java, Python and Scala programing to process the data using bigdata technologies.
- Experience in Building enterprise applications, SOAP and REST Service’s using Java/J2EE.
- Experience in working with various Big Data Distributions like Hortonworks, EMR and Databricks.
- Experience Working with streaming solution on kinesis and Spark Streaming
- Experience in migration of Data Warehouse from RDBMS to Hadoop and Cloud environments.
- Experience in working with No SQL database like HBase, etc
- Ability to optimize the usage of Hadoop to get maximum performance be it in Amazon Web Services, or In-House Cluster.
- A passion to learn new things (new Languages or new Implementations) have made me up to date with the latest trends and industry standard.
- Good experience of contributing to successful end to end analytic solutions (clarifying business objectives and hypotheses, communicating project deliverables and timelines and informing action based on findings).
- Expertise in various phases ofproject life cycles (Design, Analysis, Implementation and testing).
- Proficient in adapting to the new Work Environment and Technologies.
- Experience using integrated development environment like Eclipse, Net beans, RAD.
- Highly result oriented and pro-active, proven abilities to learn new technologies quickly and implementing them successfully in production.
- Experience in Working with various sizes of team from small to very large.
- Quick learner and self-motivated team player with excellent interpersonal skills.
- Well focused and can meet the expected deadlines on target.
- Excellent Communicational and written skills.
- Working experience in micro services using Mule run time.
- Working experience in AngularJS, Bootstrap and modern web technology stack.
TECHNICAL SKILLS
Bigdata Technologies: Spark, pyspark, Map Reduce, HBase, Hive, Pig, Sqoop, Elasticsearch/Solr, HDFS, Oozie, HCatalog, Kafka
Cloud Technologies: AWS, Databricks, S3, Lamda, EC2, SNS, AWS CLI, Cloud Watch, IAM, Kinesis, EMR, Step Functions, RDS, Data Bricks
J2EE Technologies: Servlets, JSP, JDBC, SOAP/REST Web Services, Spring
Web Servers/application servers: Apache tomcat Server, IBM WebSphere server
Web tools and languages: HTML5, CSS3, Java Script, Typescript, D3.js, High Charts, jQuery, Dojo, AngularJS, Bootstrap, npm, Web Pack
Databases: IBM DB2, Oracle Exadata, Sybase ASE, MS SQL Server 2005/2008, MySQL
Languages: Java, Python, Scala, SQL
Operating Systems: Windows 2003/2008/XP/Vista, Unix, Linux (Various Versions)
Tools: MS-Office 2003/2007/2010 , MS Access, Splunk, Data Dog, Pager Duty
Version Control: PVCS, Git Hub, Bitbucket, Jenkins
Others: Perl, Unix Scripting
IDEs: IntelliJ, PyCharm, Eclipse, RAD, Mule Any point Studio, Visual Studio
PROFESSIONAL EXPERIENCE
Confidential, Virginia
Sr Data Engineer
Responsibilities:
- Analyzing financial domain driven complex business rules and architecture and developing reusable modules in Data Pipelines using Apache Spark, Scala, Java, Python, and AWS cloud services.
- Solutioning, developing and enhancing ETL and ELT data pipelines using Apache Spark, Scala, Java, Python, Amazon S3, Amazon Lambda, Amazon EMR, Amazon EC2, Amazon SQS, Amazon SNS, Amazon ECS and Amazon RDS.
- Modeling AWS RDS database schemas.
- Monitor and support data pipelines using AWS CloudWatch, CloudTrail, PagerDuty, DataDog and Splunk.
- Providing ETL data pipeline solutions to support analytics and Machine learning teams.
- Develop and support continuous integration and continuous deployment (CI/CD) for pipelines using Git, Jenkins and cloud sentry.
Environment: JAVA/J2EE, Amazon web services, EC2, ECS, S3, EMR, RDS, SQS, SNS, Lambda, ALB, CloudTrail, CloudWatch, Cloud Formation, Hadoop, Spark, MapReduce, Scala, Java, Python, IntelliJ, Jenkins, Git, DataDog, PagerDuty, Cloud Sentry, Splunk, Snowflake, Data pipelines, postgresql, Oracle, ETL
Confidential, Berkley Heights, NJ
Big Data Developer/Data Engineer
Responsibilities:
- Developed Spark Ingestion Framework usingspark -Data Frames, Datasets for data extraction, transformation and aggregation from multiple file formats, sources for analyzing & transforming the data.
- Developed Data Quality scripts and Data reconcile using pyspark for various subject areas by comparing source and target systems
- Load current snap shot data into RDBMS ware house systems for BI and Analytics.
- Using Broadcast JOIN to improve the spark joins.
- Migrated HiveQL queries into SparkSQL to improve performance.
- Implemented POC for Real time feed with SPARK Streaming using DataBricks and segregate the data by sources.
- Extensively used Spark to analyze and process the data.
- Automate the Adhoc queries and Shared Services Processes using spark.
- Implemented POC in Apache Airflow for Scheduling the ETL Jobs.
- Using Athena to run ad-hoc queries on the S3 raw data sets.
- Used AWS lambda services by writing custom developed python scripts
- Developed serverless workflows using AWS Step Function service and automated the workflow using AWS CloudWatch
- Used PVCS/GitHub for version control and maintaining the code.
- Processing the raw XML’s using MapReduce and extracting Data into structured tables for various subject areas.
- Developed custom input formatter and record readers for XML in MapReduce Programs.
- Processing of historical large volume of XML data using big data systems.
- Designed and developed HBASE tables to maintain the state of the claims.
- Imported incremental/historical data using Sqoop from Relational Database(DB2, SQL server) to HIVE.
- Migrated the Hadoop Map Reduce in house solution to AWS.
- Developed XPATH UDF in Hive and Pig for Adhoc analysis on XML data.
- Creating Elasticsearch/ solr Collections and Loading data using Pig Scripts.
- Using Distributed Cache to store reference data and Improve performance of Map Reduce.
- Creating various views for HBASE tables and also utilizing the performance of Hive on top of HBASE.
- Developed Hive, Pig Scripts for Data Quality and Load data into Solr collections.
- Involved in Migration of Mainframe Data base to Oracle Exadata ware house and remediate the applications.
- Creating and populating Hive tables and writing hive queries for data analysis to meet the business requirements.
- Deployed and tested the application on UNIX based environments.
- Developed Claims Reporting Web Application Using Angular JS and Solr.
- Integrated Micro services using Mule.
- Implemented Spark using python and SparkSQL to cleaning, transforming the claims Notes data for text analytics using Machine learning programs.
- Worked with Data Scientist in developing of Map Reduce Streaming function using python
- Mask the sensitive PII information in claims notes datasets using DataGuise/DG Secure Tool
- Designed the Hive tables for partitioning, developed Hive SQL to load data into Hive.
- Coordinating with Business Analysts across the various countries to get user requirements and validate Cyber related Claims data sets for development.
- Developed pyspark programs to process the cyber claims data and load it into Elasticsearch/Solr collections
- Developed Rest Services API’s to consume the Solr collections data.
Environment: Hadoop, Spark, pyspark, AWS, S3, Lambda, Kinesis, Databricks,HBase, Hive, DB2, java, XML, Pig, UNIX, Scala, python, Angular Js, Solr, Oracle Exadata, Hortonworks, Git, Jenkins
Confidential
Sr Software Engineer
Responsibilities:
- Analyzing the business requirements
- Preparing of design documents and the corresponding diagrams
- Development of Front end Screens for UPD Reports & Administration modules.
- Training new associates on Front end technologies (DOJO/HTML/JS)
- Coding as per industry standards and check-lists
- Performing code review (IQA/EQA) of the other team-members
- Performing Unit-testing
- Providing Production and Pre-Production support
- Performing Root Cause analysis and providing permanent fix
- Helping business users on their queries as well on adhoc request.
- Ensuring Quality and Timeliness of Deliverable
- Preparing and delivering reports on project progress and outstanding issues
- Seeking customer feedback on-regular basis for improving the Quality of service
Environment: Java/J2EE, DOJO1.10, Angular JS, JavaScript, RESTful services, Ajax, DB2, WAS 8.7
Confidential
Software Engineer
Responsibilities:
- Development of frameworks and tools to support Surveillance reporting.
- Coordinating with compliance users and RCG to get user requirements and then working with various teams to ensure development of reports.
- Analyzing impact of data level changes on Db2 (Athena) and Sybase data warehouses.
- Working closely with various data teams to receive appropriate data as per respective report requirements.
- Performance Tuning covering SQL queries and stored procedures.
- Database change management for each release.
- Regular DB development like writing and Debugging SQL queries and Stored Procedures etc. and assisting team in SQL and other development process.
- Preparation of detailed design documents and Analysis Docs.
- Performing unit testing, System testing & UAT.
Environment: Java/J2EE, PL/SQL, DB2 LUW, Sybase ASE, Perl, UNIX
