We provide IT Staff Augmentation Services!

Big Data Engineer Resume

2.00/5 (Submit Your Rating)

Cincinnati, OH

PROFESSIONAL SUMMARY:

  • About 8+ years of professional experience in IT which includes work experience in Big Data, Hadoop ecosystem related technologies.
  • Well versed in Installation, Configuration, Supporting and Managing of Big Data and Underlying infrastructure of Hadoop Cluster.
  • Strong experience with big data processing using Hadoop technologies Map Reduce, Apache Spark, Apache Crunch, Hive, Pig and Yarn.
  • Good experience with NoSQL databases HBase, MongoDB and good understanding of Cassandra.
  • Hands on experience in using Apache SOLR/Lucene.
  • Good knowledge in streaming applications using Apache Kafka.
  • Experience in managing Hadoop clusters using Cloudera Manager tool.
  • Hands on experience with data acquisition into Hadoop cluster using Sqoop and Flume.
  • Experience in designing both time driven and data driven automated workflows using Oozie.
  • Experience with Business Intelligence tools like SAP Business Object 4.1 for creating reports.
  • Experience in analyzing data using Spark SQL, HIVEQL, PIG Latin, Spark/Scala and custom MapReduce programs in Java.
  • Good understanding of cloud configuration in Amazon web services (AWS).
  • Extending HIVE and PIG core functionality by using custom UDF’s.
  • Good experience in complete project life cycle (design, development, testing and implementation) of Client Server and Web applications.
  • Hands on Experience in application developments like Java, RDBMS and UNIX Shell Scripting.
  • Experience in Web Services using XML, HTML, Ajax, Jquery and JSON.
  • Hands on experience in J2SE, J2EE, JSP, Servlets, EJB, WebLogic, WebSphere, Tomcat, JDBC, Python and Java Script.
  • Familiarity in working with popular frameworks like Hibernate, spring, Spring Boot and MVC.
  • Good communication skills, work ethics and the ability to work in a team efficiently with good leadership skills.

TECHNICAL SKILLS:

Languages: C, C++, J2SE 1.6, 1.7, 1.8, JEE, Scala, JUnit & Shell Scripting.

Big Data Technologies: Apache Hadoop, Apache Spark, Apache Kafka, Apache Sqoop, Apache Crunch, Apache Hive, Map Reduce, Oozie, Apache NiFi and Apache Pig.

Frameworks: Spring, Spring Boot.

Web Services: RESTFUL.

Data Formats: JSON, AVRO, ORC, CSV, XML and Proto Buffer.

Data Indexing Technology: Apache SOLR.

Deployments: Pivotal Cloudy Foundry, Chef.

Integration Tools: Jenkins, Team City.

Operating Systems: Mac OS, Windows XP/ Visa/ 7

Packages & Tools: MS Office Suite (Word, Excel, PowerPoint, SharePoint, Outlook, Project)

Development Tools: Eclipse Juno, IntelliJ

Database: JDBC, MySQL, SQL Server, Oracle 10g

NoSQL Database: HBase and MongoDB.

UML Modeling Tools: Visual Paradigm for UML 10.1, Visio

BI Tools: SAP Business Objects 4.1, Information Design Tool and Web Intelligence.

PROFESSIONAL EXPERIENCE:

Confidential, Cincinnati, OH

Big Data Engineer

Responsibilities:

  • Extracted data using Sqoop Import query from multiple databases and ingest into Hive tables.
  • Nomination QA Application is developed using Scala/spark/dataframes to read data from Hive Tables on YARN Framework.
  • Evaluated and improved application performance with Spark.
  • Worked with analyst to determine and understand business requirements.
  • Implementing new dimensions into spark application upon on the business requirements.
  • Integrated various business validation & prep flag rules from campaign managers and business analyst on the data.
  • Implemented Spark Sql to update queries based on the business requirements.
  • Responsible to store processed data into MongoDB.
  • Created multi - tier java based multiple web services to read data from MongoDB.
  • Deployed Spark application and java web services in pivotal cloud foundry.
  • Design and develop the web pages as per the business requirements and user experience using Angular JS.
  • Used AUTOMIC job scheduler for scheduling multiple applications in Production.
  • Responsible for designing and developing data ingestion from Kroger using Apache NiFi/Kafka.
  • Handled Prod Deployments and provided production support for fixing the defects.
  • Agile methodology including test-driven and pair-programming concept.
  • Strong communication and analytical skills and a demonstrated ability to handle multiple tasks as well as work independently or in a team.

Environment: Apache Hadoop, Apache Spark, Apache Hive, Sqoop, Spring Boot, Pivotal Cloud Foundry, JUnit, Angular JS, IntelliJ, Maven, Team City, Automic, Apache NiFi and Git Hub.

Confidential, Cincinnati, OH

Big Data Engineer

Responsibilities:

  • Written shell scripts to extract data from Unix servers into Hadoop HDFS for long-term storage.
  • Implemented Micro Services architecture using spring boot framework.
  • Created Messaging queues using RabbitMQ to read data from HDFS to process the data.
  • Written Spark Application to implement Slowly changing dimensions (SCD Type I).
  • Created Oozie workflow in process to automate the spark application.
  • Written pig script to load processed data from HDFS into MongoDB.
  • Used MongoDB to store processed products and commodities data, which can be further down streamed into web application (Green Box/ Zoltar).
  • Deployed Spark application and java web services in pivotal cloud foundry.
  • Agile methodology including test-driven and pair-programming concept.
  • Strong communication and analytical skills and a demonstrated ability to handle multiple tasks as well as work independently or in a team.

Environment: Apache Hadoop, Apache Spark, Apache Pig, Oozie, Spring Boot, Pivotal Cloud Foundry, JUnit, IntelliJ, Maven, Team City, and Git Hub.

Confidential, Kansas City, MO

Big Data Engineer

Responsibilities:

  • Developed Crawlers java ETL framework to extract data from Cerner clients database and Ingest into HDFS & HBase for Long Term Storage.
  • Written data processing pipelines in Apache Crunch to standardize and normalize the data and store the normalized data in HBase.
  • Create ETL Pipelines using Apache Crunch to read data from HBase and calculate the KPI metrics and store the data in HDFS, and Bulk load data into HBase.
  • Experience in ETL pipeline from HDFS to Vertica.
  • Working on migrating all the Batch ETL Crunch pipelines to in-memory Spark Pipelines for quicker run times of the jobs.
  • Created Oozie workflows to manage the execution of the crunch jobs and vertica pipelines.
  • Cluster tuning to improve performance.
  • Create Integration tests to check the validity of the data being crawled.
  • Worked with other teams to determine and understand business requirement.
  • Create Hive based reports to support the application metrics (These will be used by UI teams for reports).
  • Implementing new drill to detail dimensions into the data pipeline upon on the business requirements.
  • Hands on experience with Apache SOLR for indexing HBase tables and querying the indexes.
  • Worked with multiple teams in resolving production issues.
  • Deployment automation via Chef.
  • Handled Prod Deployments and provided production support for fixing the defects.
  • Responsible for creating business layer and its underlying data foundation and connection layer using Information Design Tool.
  • Analyze the reporting requirements with solution designers and come up with the technical specifications for the Business Objects reports.
  • Created reports based on business requirements using SAP BO Web Intelligence and implemented sorting, filtering, ordering and labeling of reports.
  • Agile methodology including, test-driven and pair-programming concept.
  • Strong communication and analytical skills and a demonstrated ability to handle multiple tasks as well as work independently or in a team.

Environment: Apache Hadoop, Apache Crunch, HBase, Apache Hive, Oozie, Chops, HDFS, Apache SOLR, Splunk, SAP Business Objects 4.1, Chef, JUnit, Ruby, Eclipse, Maven, Jenkins, and Git Hub.

Confidential, Houston, TX

Hadoop Developer

Responsibilities:

  • Gathering business requirements from the Business Partners and Subject Matter Experts.
  • Installed and Configured Hadoop cluster using Amazon Web Services (AWS) for POC purposes.
  • Involved in implementing nine node CDH4 Hadoop cluster on Red hat LINUX .
  • Imported data from RDBMS to HDFS and Hive using Sqoop on regular basis.
  • Created Hive tables and worked on them using Hive QL, which will automatically invoke and run MapReduce, jobs in the backend.
  • Responsible for developing PIG Latin scripts .
  • Developed custom Map Reduce programs for data analysis and data cleaning using pig Latin scripts.
  • Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie .
  • Experience in managing and reviewing Hadoop Log files.
  • Experienced in loading and transforming large sets of structured, semi-structured and unstructured data.
  • Used Zookeeper for providing coordination services to the cluster.
  • Assisted in monitoring Hadoop cluster using Cloudera Manager.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Involved in daily SCRUM meetings to discuss the development/progress of Sprints and was active in making scrum meetings more productive.

Environment: Hadoop, Map Reduce, HDFS, Hive, Pig, Java, Hadoop distribution of Cloudera, Pig, AWS, Linux, XML, Eclipse, Oracle 10g, PL/SQL.

Confidential, Decatur, IL

Hadoop Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Responsible for Cluster maintenance, managing cluster nodes.
  • Involved in managing & review of data backups and log files.
  • Analyzed data using Hadoop components Hive and Pig.
  • Hands on experience with ETL process.
  • Involved in running Hadoop streaming jobs to process terabytes of data.
  • Experienced in loading and transforming large sets of structured, semi-structured and unstructured data Hadoop concepts.
  • Created Hive tables to store data and written Hive queries.
  • Involved in importing data from various data sources, performed transformations using Hive, Map Reduce, and loaded data into HDFS.
  • Extracted the data from Teradata into HDFS using Sqoop.
  • Exported the patterns analyzed back to Teradata using Sqoop.
  • Scheduled Oozie workflow engine to run multiple Hive and Pig jobs, which independently run with time and data availability.

Environment: Hadoop, HDFS, Hive, Pig, Sqoop, Oozie, Map Reduce, UNIX, Shell Scripting.

Confidential, Dallas, TX

Java/ J2EE/ Hadoop Developer

Responsibilities:

  • Involved in review of functional and non-functional requirements.
  • Installed and configured Hadoop MapReduce and HDFS.
  • Acquired good understanding and experience of NoSQL databases such as HBase and Cassandra.
  • Installed and configured Hive and also implemented various business requirements by writing Hive UDFs.
  • Extensively worked on user interface for few modules using HTML, JSP’s, JavaScript, Python and Ajax.
  • Generated Business Logic using servlets, Session beans and deployed them on Web logic server.
  • Created complex SQL queries and stored procedures.
  • Developed the XML schema and Web services for the data support and structures.
  • Implemented the Web service client for login verification, credit reports and applicant information using Apache Axis 2 web service.
  • Responsible for managing data coming from different sources.
  • Used Hibernate ORM framework with spring framework for data persistence and transaction management.
  • Used struts validation framework for form level validations.
  • Wrote test cases in JUnit for unit testing of classes.
  • Provided technical support for production environments resolving the issues, analyzing the defects, providing and implementing the solution defects.
  • Built and deployed Java application into multiple Unix based environments and produced both unit and functional test results along with release notes.

Environment: Hadoop, HBase, Hive, Java, Eclipse, J2EE 1.4, Struts 1.3, JSP, Servlets 2.5, WebSphere 6.1, HTML, XML, ANT 1.6, Python, JavaScript, Junit 3.8.

Confidential, Warren, NJ

Java/ J2EE Developer

Responsibilities:

  • Developed all the User Interfaces using JSP and Struts framework.
  • Writing Client Side validations using JavaScript.
  • Developed the DAO layer using the hibernate and for real time performance used the caching system for hibernate.
  • Experience in developing web services for production systems using SOAP and WSDL.
  • Developed the user interface presentation screens using HTML, XML, and CSS.
  • Experience in working with spring using AOP, IOC and JDBC template.
  • Developed the Shell scripts to trigger the Java Batch job, Sending summary email for the batch job status and processing summary.
  • Confidential -ordinate with the QA lead for development of test plan, test cases, test code and actual testing responsible for defects allocation and those defects are resolved.
  • Involved in testing and deployment of the application on Web logic Application Server during integration and QA testing phase.
  • Maintained the existing code base developed in the Struts, spring and Hibernate framework by incorporating new features and doing bug fixes.
  • Involved in configuring web.xml and struts.xml for workflow.
  • Wrote SQL queries and created DDL scripts for interacting with the Oracle database.

Environment: J2SE, J2EE, Eclipse 3.2, Spring 2.5, Hibernate 3.0, Struts 1.2, JSP, XML, Junit, Weblogic 10.3, JavaScript, Oracle 10g, HTML, AJAX, JQuery CSS.

Confidential

Software Engineer

Responsibilities:

  • Developed front-end screens using JSP, HTML and CSS.
  • Developed server side code using Struts and Servlets.
  • Developed core java classes for exceptions, utility classes, business delegate, and test cases.
  • Developed SQL queries using MySQL and established connectivity.
  • Worked with Eclipse using Maven plugin for Eclipse IDE.
  • The application was developed in Eclipse IDE and was deployed on Tomcat server.
  • Involved in scrum methodology.
  • Supported for bug fixes and functionality change.

Environment: Java, Struts 1.1, Servlets, JSP, HTML, CSS, JavaScript, Eclipse 3.2, Tomcat, Maven 2.x, MySQL, Windows, Linux.

Confidential

Intern/ Java Support Engineer

Responsibilities:

  • Extensively worked in acquiring the requirements from the business analysts and involved in all requirement clarification calls.
  • Understanding the design documents.
  • Involved in Detail level design and coding activities at offshore.
  • Involved in Code review.
  • Writing and testing the JUNIT test classes.
  • Provide support to client applications in production and other environments.
  • Working on tickets raised by the real time users and continuous interaction with end users.
  • Prepared the Technical Design Document, understanding document and test cases (UTCs and ITCs).
  • Provided Technical & Functional support to the end users during UAT & Production.
  • Continuous monitoring of application for 100% availability.

Environment: Java, Spring MVC, MIMA ORM, CVS, AQT, WebSphere, Oracle 10g and HPSM Ticketing tool.

We'd love your feedback!