We provide IT Staff Augmentation Services!

Architect Resume

0/5 (Submit Your Rating)

Tampa, FL

SUMMARY

  • 12+ years of experience in Architectural design involving large distributed systems, project management and technology implementation using BIG Data, Data warehousing, Cloud Computing, ETL, Data Visualization & RDBMS.
  • Expertise on Enterprise data platforms including Cloud & Big Data Analytics for Confidential .
  • Worked with Product owners, Business SME’s, and DI (Data Ingestion) Teams to identity requirements and consolidate enterprise data model consistent with business processes.
  • Prioritize and scale architectural efforts in close coordination with Business teams, Data Lake operational team and other stakeholders.
  • Research, evaluate and recommend tools for Enterprise Big Data Platform ingesting wide variety of data like structured, unstructured and semi structured into the Big data eco systems with processing real time data.
  • Worked on python libraries like NumPy & Pandas and created functions and data frames using Python by reading data from CSV, JSON and RDBMS.
  • Involved in migration of on - premises data to the cloud using AWS services like S3, EMR, Lambda, Athena & SQS.
  • Good experience in handling large volumes of data from various heterogeneous systems using legacy ETL tools like Ab Initio & Informatica.
  • Automated scheduling process via Airflow Scheduler and designed the code to schedule the jobs for Daily/Monthly.
  • Good exposure toSnowflake Multi-Cluster & Virtual Warehouses andinvolved in migrating existing Oracle code to Snowflake.
  • Used Splunk for monitoring the logs from various data sources and integrated all the logs for each of the source applications by building dashboards and alerts on pre-defined thresholds.
  • Good hands-on writing SQL queries and can fine tune complex queries.
  • Good communication, interpersonal skills, and strong ability to perform in a team.

TECHNICAL SKILLS

Languages: Python, SQL, CQL, PL/SQL, Unix Shell Scripting

Big Data Eco-System: Spark, HDFS, Kafka, Hive, Sqoop, Yarn

ETL: PySpark, Ab Initio, Informatica

Cloud Data Technologies: AWS, Snowflake, Splunk Cloud

Amazon Web Services: EC2, EMR, S3, Athena, SQS, CloudWatch, Glue, Lambda, RDS, Kinesis

Databases: Oracle, PostgreSQL, Cassandra, Spark-SQL, Hive, Snow-SQL

Operating Systems: Windows, Mac, Unix, Linux, Cent-OS

Scheduling Tool: Autosys, Airflow

PROFESSIONAL EXPERIENCE

Confidential, Tampa, FL

Architect

Responsibilities:

  • Designing a framework using Spark which facilitates Data Ingestion and Acquisition by creating an ETL pipeline between source to target.
  • Design and develop ETL data acquisition framework from different legacy sources of financial systems.
  • Involved in migrating data from existing Ab Initio pipelines to Spark and leveraging the system to handle multiple requests during parallel execution.
  • Performed transformations, actions, cleaning, standardization using PySpark and loaded the final dataset to HDFS.
  • Used Partitioning and bucketing in Hive for improving performance.
  • Involving in Business meetings to gather requirements, analyze/ provide tech solutions with BA.
  • Used Ab Initio for performing ETL by connecting to various sources like Flat Files, S3, HDFS and Oracle to understand the background of multifile system (MFS) functionality.
  • Working with different file formats like CSV, JSON, Avro and Parquet for processing the data.
  • Using Airflow for orchestrating the DAG’s.

Environment: Python, Ab Initio, Spark, Airflow, Bit-Bucket, Unix Shell Scripting

Confidential, San Francisco, CA

Engineering Lead

Responsibilities:

  • Design and develop ETL data acquisition framework from different legacy sources of financial systems.
  • Involved in Business meetings to gather requirements, analyze/ provide tech solutions with BA.
  • Worked with different file formats like CSV, JSON, Avro and Parquet for processing the data using PySpark and querying using Spark-SQL.
  • Automated the provisioning of the EMR cluster by Lambda function by starting up and terminating the cluster.
  • Created the SQS and SNS events/notifications for triggering the lambda functions.
  • Involved in creating ETL/ Data pipelines using Spark Framework for migrating data from existing Ab Initio pipelines.
  • Worked on creating the PySpark scripts for the data pipelines and for processing the source files.
  • Created the shell scripting for implementing the business logic and for orchestrating the PySpark scripts.
  • Created Snow pipe for continuous data load and involved in DB cloning to create separate environments.
  • Used Splunk for monitoring the logs from various data sources and integrated all the logs for each of the source applications by building dashboards and alerts based on the logs by the applications team's requirement and for fraud investigation purpose and even trigger alerts based on pre-defined thresholds.

Environment: Spark, Python, AWS, Ab Initio, Cassandra, Snowflake, Splunk, GIT, Unix, Jenkins

Confidential, San Francisco, CA

ETL Developer

Responsibilities:

  • Involve in Business meetings to gather requirements, analyze/ provide tech solutions with BA.
  • Conducted data cleansing for unstructured dataset by applyingInformatica Data Qualityto identify potential errors and improvedata integrityanddata quality.
  • Performedprofilingusing Informatica Analyst and Informatica Developer.
  • Createdlogical data object modelin Informatica developer.
  • Used Informatica Data Quality tool (Informatica Developer) to scrub, standardize and match customer address against the USPS database.
  • Design, Development and implementation ofInformatica Developer Mappingsfor data cleansing using Address validator, Labeler, Association, Parser, Expression, Filter, Router, Lookup transformations etc.
  • Used Informatica Address Doctor for global address verification across an organization.
  • Imported mapplets and mappings from Informatica developer (IDQ) to Power Center.
  • Working closely with business users to create reports/dashboards using Tableau Desktop.
  • Created Tableau worksheet which involves Schema Import, Implementing the business logic by customization.
  • Visualized transformed data by usingTableau Desktopdashboards containing histogram, trend lines, pie charts etc.

Environment: Informatica Developer/ Power Center, Tableau Desktop, Oracle, UNIX, Autosys.

Confidential

ETL Developer

Responsibilities:

  • Leading ETL system design, development and implementation efforts of business requirements from offshore.
  • Design and develop complex Ab Initio graphs/plans for data acquisition from different data sources.
  • Design and develop relational database tables/object changes for handling large volume of data.
  • Design and develop essential query scheduling engine and data validation tools for business to automate their testing process with large volume of data inputs.
  • Involved in Ab Initio Multi File system techniques, ICFF and Remote Procedure Call (RPC).
  • Worked in Ab Initio Ops-Console, continuous flows and XML components.
  • Provide performance improvement suggestions on data warehouse solutions for large scale of data.

Environment: Ab Initio, Informatica Power Center, SQL, PL/SQL, UNIX, Autosys.

Confidential

DB Developer

Responsibilities:

  • Data migration work from oracle to Postgres SQL.
  • Taken care of backend activities on Postgres DB and Oracle PLSQL.
  • Experience in developing reports using complex SQL queries.
  • Ability to write complex SQL queries for reporting purpose.
  • Involved in End to End development and deployment of one module.
  • Taking care of performance tuning activities.

Environment: Oracle, Postgres

We'd love your feedback!