Big Data Engineer Resume
4.00/5 (Submit Your Rating)
Austin, TX
EXPERIENCE SUMMARY:
- 12 Years of diverse experience in various phases of Information Technology (Development, Testing and Support). Good experience in Healthcare, Finance, Communication and Insurance domains. Currently working as Big Data Engineer.
- Overall 2 years in Big Data aspiring for a Data Engineer role in Illinois.
- Self - motivated, willing to assume responsibilities, a strong team player with great leadership qualities as well as comfortable individual with excellent communication and analytical skills.
- Self-starter and ability to adapt and learn new things quickly.
- 8 years of experience from Confidential Technology Solutions in Analytics and Information Management - ETL Programming, Big Data Programming, Data Analysis
- Good programming skills in Pig Scripting, Hive, Unix Shell Programming. Immense usage of Hadoop commands.
- Good experience in Apache Flume and Kafka.
- Currently in the process of pursuing Apache Spark Certification.
- A total of 3 years and 3 months of work experience from Convergys Information Management as Database and ETL programmer
- Strong working experience in the data analysis, SQL design and development, implementation and testing of data warehousing using extraction, transformation and loading (ETL) and SQL server and Oracle.
- Extensive experience working in Oracle PL/SQL, Teradata SQL.
- Good experience in the ETL tool Oracle Warehouse Builder.
- Worked on estimations in arriving Confidential level of effort for projects.
- Extensively used Control-M, which manages the data warehouse job, schedules.
- Good experience in performance tuning of Oracle SQL/PLSQL procedures and OWB mappings.
TECHNICAL EXPERTISE:
- Data Warehouse (ETL) and Business Intelligence
- Oracle SQL/PL-SQL
- Big Data - Hadoop, Hive QL, Pig scripting, Hbase, Kafka, Flume, Apache Spark, Yarn, Apache Cassandra
- Python Programming
- UNIX Shell Scripting
- Cognos Reports
- SQL Server
- Teradata SQL
- Oracle Warehouse Builder
- VBA - Excel Macros
- Autosys
- CONTROL - M
- PLSQL-Developer, Toad, Teradata SQL Assistant.
- ORACLE CERTIFIED ASSOCIATE (OCA) for PLSQL
- HP Quality Center
- Rally, Jira
PROFESSIONAL EXPERIENCE:
Big Data Engineer
Confidential, Austin, TX
Technology: Apache Flume, Kafka, PySpark, Spark SQL, Dataframes Cassandra, Unix Shell Programming, Yarn, Ambari, Hortonworks
Responsibilities:
- Design and Develop modules in Spark with Python using Data frames and deploy them in hadoop cluster.
- Parse data in various formats and ingest into Hive using Dataframes.
- Use Flume Services to Channelize Streaming Wifi Log Data to Kafka and HDFS
- Populate Cassandra/Hive Layers using Spark Streaming with Pyspark.
- Store Kafka offsets in Cassandra.
- Written data validation routines for realtime data coming out of Kafka.
- Ambari Cluster Monitoring - Spark, HDFS, Yarn, Zookeeper services
- Provide System Health Report to key stakeholders
- Code deployments into the cluster - Spark, Kafka, Flume
- Deploy Cassandra and Oracle DDL's
- Update configs for Flume, Kafka, Spark Jobs
- Technical Trouble Shooting of Issues in Spark, HDFS, Kafka, Flume and Cassandra clusters
- Back up, Log and System cleanup - Clean up spark history logs, spark local cache and logs
- Automate DevOps Activities - Code deployments, Yarn Cleanup, Zookeeper cleanup
ETL/Big Data Engineer
Confidential, Chicago, IL
Technology: Hive, Hbase, Teradata, Pig Scripting, Unix Shell Scripting, Hadoop
Responsibilities:
- Develop Ingestion scripts using Pig and Sqoop to ingest data into HDFS, Hive and Hbase
- Create Hive managed and external tables
- Create Unix wrapper scripts to reconcile source and target data before ingestion into HDFS
- Extensive usage of Hadoop commands
- Pull Data from RDBMS(Teradata and Oracle) Using Sqoop.
Confidential
Big Data Engineer
Technology: Oracle9i-SQL, PL/SQL,PLSQL Developer, OWB, Unix Shell Scripting
Responsibilities:
- Developing PLSQL modules for Enterprise Data Warehouse
- Deploy OWB mappings for transformations to be loaded into the enterprise data warehouse. Trouble shoot data related issues in OLAP applications and reporting.
- Use UNIX shell scripts to execute these mappings.
- Use Control-M to schedule these scripts Confidential various intervals.
- Update the code with a new version in Synergy (Version Control Repository).
- Performance optimization of SQL statements and ETL to reduce the performance overhead in OLAP databases.
- Involved in Support and Maintenance of the project after implementation in production.
- Work on Investigation and Incident requests and give the resolution with bug fix as early as possible.
- Document the DRD document after the completion of the development/fix.
- Upload the documents to the WMS system which is the central repository.
