We provide IT Staff Augmentation Services!

Bigdata Engineer Resume

3.00/5 (Submit Your Rating)

FloridA

PROFESSIONAL SUMMARY

  • Accomplished IT Professional with 7.5 years of IT experience in design and implementation of data warehousing, analytics, data integration and big data projects. This includes more than 4 years of experience in Big data Technologies.
  • Strong communication and leadership skills with sound knowledge and practical experience in data warehousing concepts. A self - motivated hardworking team player with short learning curve and the constant zeal to learn more.
  • Worked on multiple data science projects to build statistical regression model and model selection.
  • Experience in predictive analytic procedures used in supervised learning (Regression, Neural Networks, Decision trees), unsupervised learning (Clustering-k-Means and Hierarchical, PCA).
  • Experience in descriptive, exploratory, inferential, predictive modeling of the given dataset along with data modeling and recognizing key performance indicators (KPI).
  • Good experience in predictive modeling, machine learning and data mining using python and R.
  • Successfully worked and implemented multiple end-to-end Projects independently on various database applications depending on business partners requirements.
  • Extensive Experience in data modeling, data warehouse and Ralph Kimball models with Star/Snowflake Schema Designs with analysis-definition, database design, testing, and implementation and Quality process.
  • Successfully worked on leveraging multiple ETL technologies like datastage, hadoop and teradata to develop ETL applications to achieve maximum efficiency and optimum resource utilization.
  • Consulted with business partners and made recommendations to improve the effectiveness of Big Data systems, descriptive analytics systems, and prescriptive analytics systems.
  • Proficient in creating new data collection systems that optimize data management, capturing, delivery and quality
  • Working knowledge of Big Data Analytics, Hadoop ecosystems (Hadoop, Hive).
  • Experience in utilizing HIVE for working with data stored in the Hadoop file system (HDFS)
  • Sound understanding of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, and Map Reduce concepts.
  • Worked on developing big data solutions using Apache Hadoop ecosystem and tools like Hive, Spark, Sqoop, Pig, Kafka, Python, Oozie, NIFI.
  • Designed and Implemented data streaming applications to produce and consume data feeds using Kafka.
  • Successfully built and migrated multiple legacy application into hadoop environment with added efficiency and performance.
  • Extensive experience in ETL/ELT methodologies supporting Data Migration, Data Transformation and Data Cleansing of structured, unstructured and semi structured data formats using ETL tools (Informatica Power Center, DataStage).

TECHNICAL SKILLS:

BigData Tool: Hive, Pig, Oozie, Sqoop, Impala, SAS, R, Kafka

Programing: SQL, Unix Shell Scripting, Phyton

ETL Tools: IBM DataStage, Informatica

DataBases: Teradata 13.0, 14.1, MySQL, NoSQL MongoDB

Schedulers: Autosys, Oozie

Versioning: GITHUB, SVN tortoise

PROFESSIONAL EXPERIENCE:

BigData Engineer

Confidential, Florida

Responsibilities:

  • Perform End-to-end development activities, from requirement gathering and analysis, to system design, coding and testing.
  • Systematically compiled requirements and performed in-depth impact analysis
  • Develop various Teradata utilities like Mload, Bteq, FastExport etc. for required data transfer of various applications.
  • Query Tuning and Index optimization for various complex SQL queries in production.
  • Extract data to HDFS from Teradata/Oracle using Sqoop(1.4) for customer journey index application which feeds data to customer churn model.
  • Import/Export data with HDFS from Amazon Redshift(AWS) using Sqoop
  • Load data files from UNIX server to HDFS for loading into HIVE database.
  • Transform and analyze the data using HIVE (0.14) & PIG (0.12).
  • Perform performance tuning on HIVE by using concepts like partitioning and merge multiple small files etc.
  • Ingest (import/export) data to and from hdfs into rdbms using Sqoop for different kinds of file formats
  • Develop & Schedule Oozie(4.1) workflows for processing data.
  • Real time data streaming using Apache Kafka and NIFI

Data Engineer

Confidential, Jacksonville, Florida

Responsibilities:

  • Gathered requirements from client partners for application Development and implement end to end project with zero defects.
  • Implemented slowly changing Dimension logics in the mapping to effectively handle change data capture which is typical in data warehousing systems.
  • Developed automated process to perform data quality checks between two systems and generate quality reports via data stage and Hadoop technologies.
  • Leveraged multiple ETL technologies like Datastage, Hadoop and Teradata to develop ETL applications to achieve maximum efficiency and optimum resource utilization.
  • Build Data Model by analyzing the table structures to reduce Data Redundancy.
  • Involved in Several POC on Cloudera Hadoop converting small, medium, complex legacy functionality into Hadoop
  • Connected to hive backend Postgres sql and extracted critical hive metadata information and populated system dictionary tables in hive.
  • Develop & Schedule Oozie(4.1) workflows for processing data.
  • Load data files from UNIX server to HDFS for loading into HIVE database.

Teradata Applications ETL Developer

Confidential

Responsibilities:

  • Managed around 450 applications in Teradata Data Warehouse (Called The W) across different LOBs under ECIO domain in Global Technologies and operations.
  • Planning and execution of teradata hardware/Software upgrades, Disaster Recovery exercises and technology Migrations across the years and coordinated the recovery for enterprise Applications post upgrade to bring the system back to BAU.
  • Worked on application enhancements running in production environment to improve efficiency involving better run duration and resource utilization.
  • Performed ETL operations on raw data from various sources namely mainframes, Datastage and Hadoop.
  • Address user tickets and queries on the data and the applications with the business logic behind them.
  • Worked as support level 2 & 3 analyst in Datastage and Informatica platform applications. Provided 24/7 support which involves responsibilities of resolving the abends under predefined SLA’s.
  • Did performance analysis on both Datastage and target Teradata systems by analyzing for various bottlenecks and implementing indexes, collect stats and query tuning.
  • Worked on various value add automations resulting in saving both in terms of man-hours as well as the cost savings.
  • Query Tuning and Index optimization for various complex SQL queries in production.

We'd love your feedback!